Statistics on the length and linguistic complexity of bills

[This article was first published on Bommarito Consulting, LLC » r, and kindly contributed to R-bloggers]. (You can report issue about the content on this page here)
Want to share your content on R-bloggers? click here if you have a blog, or here if you don't.

  Where would you go to find out what the longest bill of the 112th Congress was by number of sections (H. R. 1473)?  How about by number of unique words (H.R. 3671)?  What about by Flesh-Kincaid reading level  (S. 475)?

  Head on over to this table of bills, updated daily for the 112th Congress, which contains the following fields:

  • Bill Name
  • Publish Date
  • Bill Title
  • Stage
  • Section Count
  • Sentence Count
  • Word (Token) Count
  • Unique Words (Tokens)
  • Unique Stem Count
  • Avg. Word Length
  • Avg. Sentence Length
  • Reading Level (Flesch-Kincaid)

I’ll be adding more automated analysis and figures over the next few weeks, but for now, here’s a morsel to get your gears turning.

To leave a comment for the author, please follow the link and comment on their blog: Bommarito Consulting, LLC » r.

R-bloggers.com offers daily e-mail updates about R news and tutorials about learning R and many other topics. Click here if you're looking to post or find an R/data-science job.
Want to share your content on R-bloggers? click here if you have a blog, or here if you don't.

Never miss an update!
Subscribe to R-bloggers to receive
e-mails with the latest R posts.
(You will not see this message again.)

Click here to close (This popup will not appear again)