## Silver Is A Weighted Coin

March 24, 2011
editorial note: there is an error in the code explained below the code. When you flip a quarter, you normally assume the coin is fair and that there is a 50% chance of getting either heads or tails. Option pricing assumes the world of trading is f...

## Generate MP3 waveforms with Ruby and R

March 24, 2011
I blame Rully for this. If it wasn’t for him I wouldn’t have been obsessed with this and spent a good few hours at night figuring it out last week. It all started when Rully mentioned that he knew how many beeps there are in the Singapore MRT (subway system) ‘doors closing’ warning. There are

## Day #11 R graphs as nodes

March 24, 2011
Today my company supervisor isn’t at work so he gave me (and the other students) a task-list. I have to check the availability of the following R scripts. Whether or not they work how they should in Knime. While doing these tasks, at the main tim...

## The Many Uses of Q-Q Plots

My last four posts have dealt with boxplots and some useful variations on that theme.  Just after I finished the series, Tal Galili, who maintains the R-bloggers website, pointed me to a variant I hadn’t seen before.  It's called a bee...

## Yeah Sure, Maybe, Well … Okay

March 23, 2011
Whoever wrote the book on statistics, probably avoided getting a proper education in literature. At least that's my null hypothesis. The cryptic and awkward presentation of probabilities common amongst the Frequentists (no, not the Latin American Socia...

## Typos sorted, at last!

March 23, 2011
After posting so many entries about typos in my books (making you wonder how there could be any text left!) and postponing their classification for so long, I decided on Saturday afternoon to collect those entries into a comprehensive pdf document that should be more useful for readers. I incidentally noticed that my book web-page

## jStat: Advanced Statistics using Javascript

March 23, 2011
While 'R' is getting enterprise ready, it's no longer the only open source option for advanced statistical programming. jStat.js is the new kid on the block.Things in favor of jStat:Based on Javascript, jQuery - future is assuredLight-weightAbility to ...

## basic ggplot2 network graphs – ver2

March 23, 2011
I posted last week a simple function to plot networks using ggplot2 package. Here is version 2. I still need to work on figuring out efficient vertex placement.Changes in version 2:-You have one of three options: use an igraph object, a matrix, or a da...

## The Popularity of Data Analysis Software (R vs SAS vs SPSS, etc.)

March 23, 2011
Robert Muenchen, the author of R for SAS and SPSS Users (A great book I’m proud to have on my shelf), has published this week an article in which he compares the popularity/market-share of many of the common statistical packages including R, SAS, SPSS and many others. The full article is available on r4stats.com at: “The Popularity of Data Analysis...

March 23, 2011
Another release (1.1.90) by Conrad Sanderson for his wonderful Armadillo templated C++ library for linear algebra appeared yesterday. Consequently, a new release 0.2.17 of RcppArmadillo, our Rcpp-based integration into R is now on CRAN mirrors. The...

## The Register profiles Revolution Analytics

March 23, 2011
Tech news site The Register has just published an in-depth profile of Revolution Analytics. It was great meeting the author Dan Olds at Revolution HQ a couple of weeks ago, and sharing with him why we think the R language is the way forward for data science: modern, applied, large-scale statistical analysis. He captures that sentiment perfectly in the...

## Graphical Display of R Package Dependencies

March 23, 2011
In some work that I am currently involved in, we have to decide which GUI engine we should use. As an obvious starter, we decided to have a look at what other people are using in their packages. While cran helpfully displays all the R packages that are available, it doesn’t (I don’t think), give

March 23, 2011
The cornerstone of your analysis and quantitative trading algorithms are data. There are lots of different ways how to do it in R (depending of what your investment instruments are). Today I am going to download data from finance.yahoo which are stock ...

## Getting into shape for the sport of data science: Screencast of talk by Jeremy Howard at Melbourne R Users

March 23, 2011
Jeremy Howard gave a talk at the Melbourne R User Group on 16th March 2011. Jeremy provided tips on how to successfully compete in data mining competitions. He showed how he combines R with other tools to build predictive models. … Continue reading →

## Applied R: Manual for the quantitative social scientist

March 23, 2011
Applied R for the quantitative social scientist is a manual on R written specifically as an introduction for the quantitative social scientist. To my opinion, R-Project is a magnificent statistical program, ready to be accepted and implemented in the social sciences. The flexibility of this program and the way data are handled gives the user a sense of closeness...

## sab-R-metrics Sidetrack: Bubble Plots

March 22, 2011
While I had mentioned in my last post that I will cover logistic regression in my next post, I decided that a quick interlude in working with bubble plots would be fun. Bubble plots have become pretty popular recently, especially with all of the Visualization Challenges I've seen around the internet (by the way, I...

## R again in Google Summer of Code

March 22, 2011
I'm a big fan of the Google Summer of Code.  It brings great projects together with a learning opportunity for students.  Once again the R Project was selected to be part of the Google Summer of Code in 2011.  Some other notable mathemat...

## Where the heck has JD been?

March 22, 2011
It’s been pointed out to me that I haven’t had any blog posts in a while. It’s true. I’m fairly slack. But in the last few months I’ve changed jobs (same firm, new role), written an R abstraction on top of Hadoop, been to China, and managed to stay married. While that sounds pretty awesome,

## Code: extended model support for mtable

March 22, 2011
$Code: extended model support for mtable$

I finally got around to organizing and packaging my complete set of extended model support for mtable in Martin Elff’s memisc library. Here is a list of the models supported: coxph, survreg – Cox proportional hazards models and parametric survival … Continue reading →

## More on R-Studio

March 22, 2011
Here's a link to keyboard shortcuts for the RStudio IDE. RStudio has replaced EMacs, Aquamacs, Tinn-R, Bluefish, and even Komodo Edit as my preferred R IDE/editor.http://gettinggeneticsdone.blogspot.com/2011/03/rstudio-keyboard-shortcut-reference-pdf.h...

## Analysis: R growth continues in popularity of data analysis software

March 22, 2011
Bob Muenchen, author of R for SAS and SPSS Users and co-author of R for Stata Users, has updated his in-depth analysis of the popularity of data analysis software. Determining "popularity" for software is a tricky task, but this analysis looks at several different metrics: mailing list traffic, blogs, search volumes, job listings and other such indirect methods, as...

## One bicycle for two

March 22, 2011
$One bicycle for two$

Robin showed me a mathematical puzzle today that reminded me of a story my grand-father used to tell. When he was young, he and his cousin were working in the same place and on Sundays they used to visit my great-grand-mother in another village. However, they only had one bicycle between them, so they would

## Free and Easy Currency Monitor in R

March 22, 2011
Certainly not the best way to keep up with currencies, but the increasingly important job of monitoring currencies can be free and easy in R using Federal Reserve FRED data.  Here is a template that can be adjusted to your favorite currencies with...

## A Short Side-by-side Comparison of the R and NumPy Array Types

March 22, 2011
Feature NumPy R contiguous (virtual) memory ✔ ✔ 'view' memory model ✔ ✘ subset-assignment ✔ ✔ vectorized operations ✔ ✔ memory-mapping ✔ ✘* broadcasting rules ✔ ✔ index arrays ✔ ✔ This comparison is current as of R 2.13.0, NumPy version 1.4.1, and other web resources to date. Because this post was motivated by a

## Visualizing Missing Data

March 22, 2011
There are several graphics available for visualizing missing data. The following graphic was inspired by many sources. However, I wanted a version using ggplot2. What is visualized here is the percent missing for each variable in the PISA data across countries. The code will be available as part of the multilevelPSA package I am currently

## data.table: an R package everyone should use

March 22, 2011
I’m not sure how I missed this package, but I am sure glad I’ve found it. The data.table package for R provides something of a reconceptualization of the standard data.frame object. Though it remains (mostly) compatible with data.frame. The advantage … Continue reading →