Packt+ | Advance your knowledge in tech

You're reading from scikit-learn Cookbook , Second Edition Over 80 recipes for machine learning in Python with scikit-learn

Product type Paperback

Published in Nov 2017

Publisher Packt

ISBN-13 9781787286382

Length 374 pages

Edition 2nd Edition

Languages

Python

Tools

Scikit-learn

Concepts

Machine Learning

Author (1):

Trent Hauck

View More author details

Table of Contents (19) Chapters

Title Page

Credits

About the Authors

About the Reviewer

www.PacktPub.com

Customer Feedback

Preface

1. High-Performance Machine Learning – NumPy FREE CHAPTER

2. Pre-Model Workflow and Pre-Processing

3. Dimensionality Reduction

4. Linear Models with scikit-learn

5. Linear Models – Logistic Regression

6. Building Models with Distance Metrics

7. Cross-Validation and Post-Model Workflow

8. Support Vector Machines

9. Tree Algorithms and Ensembles

10. Text and Multiclass Classification with scikit-learn

11. Neural Networks

12. Create a Simple Estimator

Machine learning with logistic regression

You are familiar with the steps of training and testing a classifier. With logistic regression, we will do the following:

Load data into feature and target arrays, X and y, respectively
Split the data into training and testing sets
Train the logistic regression classifier on the training set
Test the performance of the classifier on the test set

Getting ready

Define X, y – the feature and target arrays

Let's start predicting with scikit-learn's logistic regression. Perform the necessary imports and set the input variables X and the target variable y:

import numpy as np
import pandas as pd

X = all_data[feature_names]
y = all_data['target']

How to do it...

Provide training and testing sets

Import train_test_split to create testing and training sets for both X and y: the inputs and target. Note the stratify=y, which stratifies the categorical variable y. This means that there are the same proportions of zeros and ones in both y_train and y_test: