jsfenfen/whatwordwhere
Tooling to extract data from scanned paper forms OCR-ed by Tesseract using the HOCR standard.
Discovered public repositories for jsfenfen in the GitHub catalog.
Tooling to extract data from scanned paper forms OCR-ed by Tesseract using the HOCR standard.
muck with sunlight house disbursement csvs
another bucket of scripts for grabbing the fec's ftp data etc for django + postgres
Parse XBRL filings from the SEC's EDGAR in Python
Members of the United States Congress, 1789-Present, in YAML, as well as committees, presidents, and vice presidents.
A Ruby parser for electronic candidate, PAC and party campaign filings from the Federal Election Commission.
Public repository.
Data from the census bureau's "easy stats" site--the first available on the 113th Congress.
Test open refine reconciliation service to match legislators names
Data from the Plum Book, published by the GPO every 4 years
US state metadata and other incredible fun stuff
Legacy export of ACS processing from 2008 3-year ACS for R and PostgreSQL
Add some fuzzy string match operations to postgreSQL
Noodle with document cloud
like inspectdb, but for files