Logistic Regression
The logistic regression model is also one of the most popular models in machine learning for performing classification ( predicting the values of a qualitative variable based on predictors ). It is a parametric technique - the model must find the best parameters based on the data -.
This technique is used to adjust the relationship between a qualitative variable ( dependent variable ) and a set of predictors - independent variables - which must be quantitative variables or qualitative variables transformed into quantitative variables ( one-hot encoding or discrete variable digitization ).
The operation of logistic regression is almost identical to linear regression except that it uses a sigmoid function. Like linear regression, we assume that and are dependent, meaning that knowing the values of improves the knowledge of the values of . Therefore, there is a correlation between and - as a reminder : the more a variable is correlated ( positively or negatively ) with the variable , the more important it is for our model because we say it is « discriminant ». – see ( correlation ).
Logistic regression does not directly predict a qualitative value but a probability that a new record belongs to a class.
Linear vs Logistic
When we talk about regression, it refers to any process that tends to find relationships between variables.
Linear regression seeks to establish a linear relationship between the dependent variable (y) and the explanatory or independent variables () ( predictors ). Example of simple linear regression : if the car’s age increases by 1 year, the price will be impacted by ( according to the coefficient and the origin of the line ).
Logistic regression also seeks to establish a relationship between the dependent variable (y) and the explanatory or independent variables () ( predictors ) but uses a logistic function ( logit ) to obtain a value between and , which is the probability that a new record belongs to a class. For example, if the car’s age increases by 1 year, the probability that it breaks down will increase or decrease depending on the coefficient associated with age in the model.
One uses a linear function while the other uses a sigmoid function which, through an underlying linear function, models a probability.

Why Not Use Linear Regression ?
Let’s take the example of a classification with two values : or . We have a distribution of values for both cases. If we use linear regression and draw the line, we could, in the example illustrated below on the left, consider that linear regression does a correct job, assuming we define a threshold of 0.5 to say that any value equal to or greater than equals ( class 1 ) and any value below equals ( class ). We see in the left graph that all values to the right of the perpendicular to have a threshold greater than or equal to and thus .
If suddenly, we have a record where the value of is much higher ( right graph – purple dot ), in this case, the slope of the line is modified, and we see that the model would predict some values greater than or equal to the threshold as belonging to class .

Therefore, we need to use a function that allows us to obtain a curve with the particularity of applying a linear form in a sub-function when values increase exponentially, but that would be flat at the beginning and end so that the values always remain between and . Specifically, it is an S-shaped function to ensure that possible values are between and .

Sigmoid Function
The sigmoid function - logistic function - is an S-shaped mathematical function that transforms any value into a number between and .
or
represents the linear sub-function. Just like linear regression, the logistic regression algorithm has training data, i.e. and , and based on this information, it must identify the parameters through a cost function and the gradient descent algorithm to find the local optimum. Once the parameters are identified, the sigmoid model outputs a value between 0 and 1 – a probability – which is transformed into a class based on a threshold definition ( usually ).
When rewritten in detail, the formula for gives us this :
and therefore the formula for the sigmoid function :
or
If we break down the formula :
- is the output of the sigmoid function, which gives us a value between and ;
- is the input to the function - which contains the parameters identified by the gradient descent algorithm based on the training data -. also requires new values of to predict a probability ;
- is Euler’s constant ( approximate value of ) and plays a crucial role in obtaining the S-shape ;
No matter the value of , the negative exponent of Euler's constant in the denominator will ensure that :
- if is a very large positive value , the sigmoid function will approach a maximum of ;
- if is a very large negative value , the sigmoid function will approach a minimum of ;
Sigmoid function in Python
def sigmoid(z):
g = 1/(1+np.exp(-z))
return g
# (z = np.dot(X[i],w)+b) - define later in the cost function
Decision Threshold
Logistic regression allows us to output a value between and from the function , and we, as data scientists, must define the threshold for class or class . In the literature, the threshold of is often mentioned, such that if the prediction value is , then , and if the prediction value is , then .
Linear Threshold
In the case of a linear decision based on two predictors ( ), we could represent the data as follows (diagram below) for values where , which corresponds to
To simplify the example, let's define 1 for and , and -3 for .
The linear decision threshold in this case would correspond to the moment when : as this threshold would be neutral in defining whether (red cross) or (blue circle).
In our case, since and are both 1 and is -3, we can rewrite the equation as follows : . Therefore, .

The formula for the linear decision threshold corresponds to the initial formula for logistic regression :
or
Non-Linear Threshold
As we discussed for polynomial regression in the chapter on linear regression, we may encounter cases in logistic regression where the data separation is non-linear.

In the case of a non-linear threshold, we will use polynomial features by adapting the formula as follows ( for two predictors and ) :
or
Other degrees of polynomial forms are, of course, tested within the algorithm to find the best way to predict the information. For example, in the following cases :

The equation would take the following form :
import numpy as np
import pandas as pd
from sklearn.preprocessing import PolynomialFeatures, StandardScaler
from sklearn.model_selection import train_test_split
import math, copy
from sklearn.metrics import confusion_matrix
# Data split
data = pd.read_csv('data.csv')
# Variables split
x = data[['x_1', 'x_2', 'x_3', 'x_4', 'x_5', '...', 'x_n']].values
y = data['y'].values
# Generation of 3rd Degree Polynomial Terms
poly = PolynomialFeatures(degree=3)
x_poly = poly.fit_transform(x)
# Standardization of Polynomial Data
scaler_x = StandardScaler()
x_poly = scaler_x.fit_transform(x_poly)
# Splitting into Training and Test Sets
x_train, x_test, y_train, y_test = train_test_split(x_poly, y, test_size=0.4, random_state=42)
Cost Function
The shape of the cost function in linear regression is convex, which is optimal for the gradient descent algorithm, whose goal is - through iteration -to find the optimal values for and .
In linear regression, we use the following formula for the cost function : The objective is then to find the parameters that minimize this cost, .
The problem with logistic regression is that the shape of its cost function is non-convex, meaning that gradient descent can find several local optima that may not correspond to the true minimum and can get stuck in convergence.

Logistic regression uses a « transformed » cost function from linear regression to make the cost function convex again, allowing it to converge towards a local optimum.
Loss Function
To adapt the cost function to logistic regression, let’s focus on the concept of loss, which can be isolated as part of the cost function , shown in blue in the formula below :
This loss function for all records can be rewritten as :
If we visualize by considering on the x-axis ( abscissa ), knowing that is the result of logistic regression ( a value between and ), it would resemble the green curve in the graph below ; and if we visualize , we would obtain the blue curve.
The intersection of the two curves on the x-axis corresponds to the value . The part of the function that concerns the result between and is the top left section, outlined in red.

If we zoom in on this section and assume for the loss function and our model predicts :
- If the model predicts , the loss is ;
- If the model predicts , the loss is ( natural logarithm ) ;
- If the model predicts , the loss is ;

What we observe is that the algorithm will aim to reduce the loss and be as accurate as possible because predictions where but result in significant loss.

In the second case (), the algorithm will also aim to reduce the loss, and predictions where but result in significant loss.
Therefore, transforming the cost function :
Allows us to obtain a convex form and use gradient descent to find the local optimum.
Of course, the cost function applies to the entire dataset and will correspond to the sum of losses divided by :
The gradient descent algorithm will therefore try to find the parameters that minimize the cost function.
The loss function :
Can be simplified as follows :
It is simplified because it combines both cases ( or ) and automatically cancels one case if the other is assumed.
If and we substitute the values of into the formula :
The right-hand side is canceled because , and thus :
As a result, we get :
If and we substitute the values of into the formula: