Public GitHub Catalog
Discovered Sep 29, 2026
GregorySenay / wikipedia-extractor
Extracts and cleans text from Wikipedia database dump and stores output in a number of files of similar size in a given directory. This is a mirror of the script by Giuseppe Attardi.
About this discovery
This repository is cataloged as part of our automated global GitHub synchronization. Full telemetry, velocity snapshots, and code summaries are scheduled for continuous enrichment.
#15463757GitHub System ID
GregorySenayOrganization / User
PublicVisibility
ActiveCatalog Status