XERJ is the new way for AI to search data. Its autoindex capability activates agents to know your data without the token waste of grep and sed. One command indexes code, docs, logs and PDFs for search, RAG, security audits and agent memory, using 40x fewer tokens than grep. Elasticsearch compatible, so existing clients just work.
Top BM25 GitHub Repositories & Tools (2026)
Discover the most starred and trending open source tools tagged with #bm25.
Git-native persistent memory for AI coding agents. Implements Google OKF v0.2 with sub-300µs in-memory BM25 search, embedded MCP server, and progressive disclosure. Slashes token bloat by 80% with zero external databases or dependencies. Built in pure Go.
One Postgres for your application data, full-text search, vector retrieval, and aggregations. Home of the pg_search extension.
Open-source search database for full-text, vector, and hybrid search with real-time indexing and SQL.
A queryable second brain over your scattered notes and docs - hybrid retrieval (vector + BM25), section-level citations, and an MCP server so AI agents can use it. ~300 lines, no LangChain.