Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
A collection of datasets and tasks for legal machine learning
| Date | Stars |
|---|---|
| 2026-07-31 | 441 |
| 2026-08-02 | 441 |
| 2026-08-06 | 443 |
Today
+2 stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# Datasets for Machine Learning in Law This is a collection of pointers to datasets/tasks/benchmarks pertaining to the intersection of machine learning and law. If you want to add a dataset, create a pull request (following the format below). Alternatively, email [Neel Guha]([email protected]). Neel Guha --- ### [A Corpus of eRulemaking User Comments for Measuring Evaluability of Arguments](https://facultystaff.richmond.edu/~jpark/papers/jpark_lrec18.pdf) >eRulemaking is a means for government agencies to directly reach citizens to solicit their opinions and experiences regarding newly proposed rules. The effort, however, is partly hampered by citizens’ comments that lack reasoning and evidence, which are largely ignored since government agencies are unable to evaluate the validity and strength. We present Cornell eRulemaking Corpus – CDCP, an argument mining corpus annotated with argumentative structure information capturing the evaluability of arguments. The corpus consists of 731 user comments on Consumer Debt Collection Practices (CDCP) rule by the Consumer Financial Protection Bureau (CFPB); the resulting dataset contains 4931 elementary unit and 1221 support relation annotations. It is a resource for building argument mining systems that can not only extract arguments from unstructured text, but also identify what additional information is necessary for readers to understand and evaluate a given argument. Immediate applications include providing real-time feedback to commenters, specifying which types of support for which propositions can be added to construct better-formed arguments. ### [A Dataset for Statutory Reasoning in Tax Law Entailment and Question Answering](https://ceur-ws.org/Vol-2645/paper5.pdf) >Legislation can be viewed as a body of prescriptive rules expressed in natural language. The application of legislation to facts of a case we refer to as statutory reasoning, where those facts are also expressed in natural language. Computational statutory reasoning is distinct from most existing work in machine reading, in that much of the information needed for deciding a case is declared exactly once (a law), while the information needed in much of machine reading tends to be learned through distributional language statistics. To investigate the performance of natural language understanding approaches on statutory reasoning, we introduce a dataset, together with a legal-domain text corpus. Straightforward application of machine reading models exhibits low out-of-the-box performance on our questions, whether or not they have been fine-tuned to the legal domain. We contrast this with a hand-constructed Prolog-based system, designed to fully solve the task. These experiments support a discussion of the challenges facing statutory reasoning moving forward, which we argue is an interesting real-world task that can motivate the development of models able to utilize prescriptive rules specified in natural language. ### [ACL/CoLing Dataset](https://usableprivacy.org/data) >We created a corpus of 1,010 privacy policies from the top websites ranked on Alexa.com. The privacy policies in the dataset were retrieved in December 2013 and January 2014. ### [Affidavit Verification] (https://huggingface.co/datasets/datuk2/affi2) >From: [Chris Kwan](https://huggingface.co/datuk2/datasets): "This may be redundant as the latest LLM seem to know and able to identify statements like hearsay and so on. This data-set was build based on personal experiences of how one identify statements that are liable to be challenged by opposition. It did not go far obviously since ones person effort is limited. So I am hoping it can draw on many others so a much useful data-set can be build." ### [APP-350 Dataset](https://usableprivacy.org/data) >The APP-350 Corpus consists of 350 Android app privacy policies annotated with privacy practices (i.e., behavior that can have privacy implications). ### [AsyLex: A Dataset for Legal Language
Excerpt of 46,138 characters
Read on GitHubWould you bet a product on this? Bounded 0–100 and slow moving.
matched fp:a10985c1f2f6237c, name:datasets, desc:datasets