generative AI for data analysis

Unlocking 20M Indian court cases with structured legal data and AI-ready embeddings.

Introducing a comprehensive dataset of over 20 million Indian legal documents, meticulously structured for legal NLP applications.

3 min readMachine Learning

For two years, someone has been doing the unglamorous, painstaking work of turning India's legal chaos into structure. The result is a dataset of over 20 million court cases, complete with metadata, citation graphs, and AI-ready embeddings. This is not a flashy demo or a press release. It is the kind of foundational infrastructure that legal tech has been quietly waiting for, and the person behind it deserves attention.

What matters most here is the citation graph. Knowing that Case A followed, distinguished, or overruled Case B across millions of judgments is not just a convenience. It is the difference between guessing how legal precedent works and actually tracing its path. For anyone building retrieval-augmented generation systems, this is a goldmine. If your retriever can't surface the case that a judgment explicitly relies on, then your system is guessing, not reasoning. This dataset gives you a way to test that, with ground truth built into the relationships themselves. That is rare. That is valuable.

The practical applications go beyond research papers. Legal outcome prediction, influence analysis, and graph neural networks all become more feasible when the data is already structured. The fact that the metadata pipeline also identifies judges, advocates, sections, and acts from unstructured text means this could serve as training data for legal NER models. And for those working on low-resource Indian languages, the bilingual pairs from the translation service offer a different register than typical news or conversational corpora. Legal language is formal, precise, and domain-specific. Having that in a training set is a distinct advantage.

There are honest limitations here, and the author acknowledges them. English is the primary language for most judgments, with regional language data coming from translations rather than originals. Metadata accuracy varies by court, and the citation classification is estimated at 90-95% precision. That is fine. No dataset is perfect, and pretending otherwise would undermine the credibility of the work. What matters is that the foundation is solid, the scope is massive, and the door is open for feedback.

If you are working on legal NLP, graph-based analysis, or Indian language models, this is worth your time. The data is available via API and bulk export, and it is public domain. The dataset's creator is asking what would be most useful. That is the right question to ask, and it signals a willingness to build with the community rather than just for it. The opportunity is here. The question is who will use it.

From Machine Learning

been working on structuring India's legal corpus for the past 2 years and wanted to share what I've built and hear from people working on legal NLP or low-resource Indian language models.

dataset is 20M+ Indian court cases from the Supreme Court, all 25 High Courts, and 14 Tribunals. each case has structured metadata (court, bench, date, parties, judges, sections cited, acts referenced, case type). there's a citation graph across the full corpus where I've classified relationships as followed, distinguished, overruled, or mentioned.

Read the original at Machine Learning