# 2220 search results for "Regression"

## Visualizing Bootrapped Stepwise Regression in R using Plotly

May 29, 2016
By

We all have used stepwise regression at some point. Stepwise regression is known to be sensitive to initial inputs. One way to mitigate this sensitivity is to repeatedly run stepwise regression on bootstrap samples. R has a nice package called bootStepAIC() which (from its description) “Implements a Bootstrap procedure to investigate the variability of model

## End to end Logistic Regression in R

May 29, 2016
By

Logistic regression, or logit regression is a regression model where the dependent variable is categorical. I have provided code below to perform end-to-end logistic regression in R including data preprocessing, training and evaluation. The dataset used can be downloaded from here.

## Principal Components Regression in R: Part 2

May 24, 2016
By

by John Mount Ph. D. Data Scientist at Win-Vector LLC In part 2 of her series on Principal Components Regression Dr. Nina Zumel illustrates so-called y-aware techniques. These often neglected methods use the fact that for predictive modeling problems we know the dependent variable, outcome or y, so we can use this during data preparation in addition to using...

## Principal Components Regression, Pt. 2: Y-Aware Methods

May 23, 2016
By

In our previous note, we discussed some problems that can arise when using standard principal components analysis (specifically, principal components regression) to model the relationship between independent (x) and dependent (y) variables. In this note, we present some dimensionality reduction techniques that alleviate some of those problems, in particular what we call Y-Aware Principal Components … Continue reading...

## Visual contrast of two robust regression methods

Robust regression For training purposes, I was looking for a way to illustrate some of the different properties of two different robust estimation methods for linear regression models. The two methods I’m looking at are: least trimmed squares, implemented as the default option in lqs() a Huber M-estimator, implemented as the default option in rlm() Both functions...

## Principal Components Regression in R, an operational tutorial

May 17, 2016
By

John Mount Ph. D. Data Scientist at Win-Vector LLC Win-Vector LLC's Dr. Nina Zumel has just started a two part series on Principal Components Regression that we think is well worth your time. You can read her article here. Principal Components Regression (PCR) is the use of Principal Components Analysis (PCA) as a dimension reduction step prior to linear...

## Principal Components Regression, Pt.1: The Standard Method

May 16, 2016
By

In this note, we discuss principal components regression and some of the issues with it: The need for scaling. The need for pruning. The lack of “y-awareness” of the standard dimensionality reduction step. The purpose of this article is to set the stage for presenting dimensionality reduction techniques appropriate for predictive modeling, such as y-aware … Continue reading...

## Manipulate(d) Regression!

May 5, 2016
By

The R package ‘manipulate’ can be used to create interactive plots in RStudio. Though not as versatile as the ‘shiny’ package, ‘manipulate’ can be used to quickly add interactive elements to standard R plots. This can prove useful for demonstrating statistical concepts, especially to a non-statistician audience. The R code at the end of this

## How long could it take to run a regression

April 6, 2016
By
$n$

This afternoon, while I was discussing with Montserrat (aka @mguillen_estany) we were wondering how long it might take to run a regression model. More specifically, how long it might take if we use a Bayesian approach. My guess was that the time should probably be linear in , the number of observations. But I thought I would be good to check. Let...

## Assessing significance of slopes in regression models with interaction

March 17, 2016
By

This is a pretty short post on an issue that popped at some point in the past, at that time I found a way around it but as it arose again recently I decided to go through it. The issue I had was that when modeling an interaction between a continuous (say temperature) and a Related Post