Sampling and the Analysis of Big Data
After my last post, I came across a few articles supporting the opinion that if you have a good reason to take random samples from a “big” dataset, you’re not committing some kind of sin:
Big Data Blasphemy: Why Sample?
To Sample or Not to Sample… Does it Even Matter?
The moral of the story is that you can sample from “big data” so long as the analysis you’re doing doesn’t require some part of the data that will be excluded as part of the sampling process (an exampl being the top or bottom so many records based on some criterion).
offers daily e-mail updates
news and tutorials
on topics such as: visualization (ggplot2
), programming (RStudio
, Web Scraping
) statistics (regression
, time series
) and more...
If you got this far, why not subscribe for updates
from the site? Choose your flavor: e-mail
, or facebook