A classifier takes a set of measurements and says which class it belongs to. Give it the height, weight and shoe size of a person and it can tell you which label that is in your training data.
You never write the rules yourself. The algorithm finds the boundary from the training data.
A classifier takes a set of measurements and says which class it belongs to. Give it the height, weight and shoe size of a person and it can tell you whether that is a male or a female label in your training data.
You never write the rules yourself. The algorithm finds the boundary from the training data.
Training data goes into the algorithm, it works out where the classes separate, and then you can predict with it.
Machine Learning Classification
In the example below we predict if it’s a male or female given vector data.
We start with training data. In this example we have a set of vectors (height, weight, shoe size) and the class this vector belongs to:#{height, weights, shoe size}
X = [[190,70,44],[166,65,45],[190,90,47],[175,64,39],[171,75,40],[177,80,42],[160,60,38],[144,54,37]]
Y = ['male','male','male','male','female','female','female','female']
Define a vector for your prediction in the same format (height, weight, size). If you want, you can also get this from console input:P = [[190,80,46]]
Then we fit the training data and predict in this style:c = Classifier()
c = c.fit(X,Y)
print "\nPrediction : " + str(c.predict(P))

That gives us this code:from sklearn.tree import DecisionTreeClassifier
from sklearn.neighbors import KNeighborsClassifier
from sklearn.neural_network import MLPClassifier
from sklearn.ensemble import RandomForestClassifier
#{height, weights, shoe size}
X = [[190,70,44],[166,65,45],[190,90,47],[175,64,39],[171,75,40],[177,80,42],[160,60,38],[144,54,37]]
Y = ['male','male','male','male','female','female','female','female']
#Predict for this vector (height, wieghts, shoe size)
P = [[190,80,46]]
#{Decision Tree Model}
clf = DecisionTreeClassifier()
clf = clf.fit(X,Y)
print "\n1) Using Decision Tree Prediction is " + str(clf.predict(P))
#{K Neighbors Classifier}
knn = KNeighborsClassifier()
knn.fit(X,Y)
print "2) Using K Neighbors Classifier Prediction is " + str(knn.predict(P))
#{using MLPClassifier}
mlpc = MLPClassifier()
mlpc.fit(X,Y)
print "3) Using MLPC Classifier Prediction is " + str(mlpc.predict(P))
#{using MLPClassifier}
rfor = RandomForestClassifier()
rfor.fit(X,Y)
print "4) Using RandomForestClassifier Prediction is " + str(rfor.predict(P)) +"\n"
How it learns
- You give it training data, one row of numbers per sample, plus the class for each row.
- You call
fit()and the algorithm works out where the classes separate. - You call
predict()with a new row and get the class back. - That separating surface is called the decision boundary.
- Accuracy is the share of predictions that were right.
score()gives it to you.
Different classifiers draw the boundary in different shapes. A decision tree draws steps, k-nearest neighbors follows the shape of the data and an SVM draws a straight line between the classes.
Want more practice on this? Practise this on PyChallenge, there are short browser exercises you can run right after reading.
