Built something? We create video reels & spotlights for GitHub projects.Promote your project →
Catalog / scraping-xx / corpusmaker
Public GitHub Catalog Discovered Sep 27, 2026

scraping-xx / corpusmaker

clojure utilities to build training corpora for machine learning / NLP out of public wikimedia dumps: status - partially stalled - will probably be reworked as cascalog scripts -- this project is in stalled mode right now: the pignlproc project is likely to replace it due to licensing constraints for future integration in Apache projects

View repository on GitHub ↗ View creator profile Browse directory

About this discovery

This repository is cataloged as part of our automated global GitHub synchronization. Full telemetry, velocity snapshots, and code summaries are scheduled for continuous enrichment.

#9453765GitHub System ID
scraping-xxOrganization / User
PublicVisibility
ActiveCatalog Status

More from scraping-xx

FOR MAINTAINERS

Built something? Put it in front of millions of developers.

We make a short reel about your project and post it across YouTube, Instagram, Threads, and X. Send a link, we do the rest.