Top AI Repos โ open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Python suite to construct benchmark machine learning datasets from the MIMIC-III ๐ clinical database.
| Date | Stars |
|---|---|
| 2026-07-31 | 891 |
| 2026-08-06 | 891 |
| 2026-08-07 | 891 |
| 2026-08-15 | 891 |
| 2026-08-16 | 891 |
| 2026-08-18 | 891 |
| 2026-08-20 | 890 |
| 2026-08-31 | 889 |
| 2026-09-10 | 890 |
| 2026-09-19 | 891 |
| 2026-09-20 | 892 |
Today
+1 stars today
This week
+2 stars this week
This month
+2 stars this month
Momentum
0.0
growth rate 0.22%/day
MIMIC-III Benchmarks
=========================
[](https://gitter.im/YerevaNN/mimic3-benchmarks?utm_source=badge&utm_medium=badge&utm_campaign=pr-badge&utm_content=badge)
Python suite to construct benchmark machine learning datasets from the MIMIC-III clinical database. Currently, the benchmark datasets cover four key inpatient clinical prediction tasks that map onto core machine learning problems: prediction of mortality from early admission data (classification), real-time detection of decompensation (time series classification), forecasting length of stay (regression), and phenotype classification (multilabel sequence classification).
## News
* 2018 December 28: The second draft of the paper is released on [arXiv](https://arxiv.org/abs/1703.07771).
* 2017 December 8: This work was presented as a spotlight presentation at NIPS 2017 [Machine Learning for Health Workshop](https://ml4health.github.io/2017/).
* 2017 March 23: We are pleased to announce the first official release of these benchmarks. We expect to release a revision within the coming months that will add at least ~50 additional input variables. We are likewise pleased to announce that the manuscript associated with these benchmarks is now [available on arXiv](https://arxiv.org/abs/1703.07771).
## Citation
If you use this code or these benchmarks in your research, please cite the following publication.
```
@article{Harutyunyan2019,
author={Harutyunyan, Hrayr and Khachatrian, Hrant and Kale, David C. and Ver Steeg, Greg and Galstyan, Aram},
title={Multitask learning and benchmarking with clinical time series data},
journal={Scientific Data},
year={2019},
volume={6},
number={1},
pages={96},
issn={2052-4463},
doi={10.1038/s41597-019-0103-9},
url={https://doi.org/10.1038/s41597-019-0103-9}
}
```
**Please be sure also to cite the original [MIMIC-III paper](http://www.nature.com/articles/sdata201635).**
## Motivation
Despite rapid growth in research that applies machine learning to clinical data, progress in the field appears far less dramatic than in other applications of machine learning. In image recognition, for example, the winning error rates in the [ImageNet Large Scale Visual Recognition Challenge](http://image-net.org/challenges/LSVRC/) (ILSVRC) plummeted almost 90% from 2010 (0.2819) to 2016 (0.02991).
There are many reasonable explanations for this discrepancy: clinical data sets are [inherently noisy and uncertain](http://www-scf.usc.edu/~dkale/papers/marlin-ihi2012-ehr_clustering.pdf) and often small relative to their complexity, and for many problems of interest, [ground truth labels for training and evaluation are unavailable](https://academic.oup.com/jamia/article-abstract/23/6/1166/2399304/Learning-statistical-models-of-phenotypes-using?redirectedFrom=PDF).
However, there is another, simpler explanation: practical progress has been difficult to measure due to the absence of community benchmarks like ImageNet. Such benchmarks play an important role in accelerating progress in machine learning research. For one, they focus the community on specific problems and stoke ongoing debate about what those problems should be. They also reduce the startup overhead for researchers moving into a new area. Finally and perhaps most important, benchmarks facilitate reproducibility and direct comparison of competing ideas.
Here we present four public benchmarks for machine learning researchers interested in health care, built using data from the publicly available Medical Information Mart for Intensive Care (MIMIC-III) database ([paper](http://www.nature.com/articles/sdata201635), [website](http://mimic.physionet.org)). Our four clinical prediction tasks are critical care variants of four opportunities to transform health care using in "big clinical data" as described in [Bates, et al, 2014](http://content.healthaffairs.orExcerpt of 16,306 characters
Read on GitHub203
42
16
2
1
The Gitter Badger ยท Gitter
1
1
1
1
1
Would you bet a product on this? Bounded 0โ100 and slow moving.
matched fp:b49b912760390125, llm:Repository topics and description: 'benchmark, clinical-data, deep-learning, machine-learning' and README: 'Python suite to construct benchmark machine learning datasets from the MIMIC-III clinical database... benchmark datasets cover ... classification, time series classification, regression, multilabel sequence classification.'
matched fp:b49b912760390125, llm:Repository topics and description: 'benchmark, clinical-data, deep-learning, machine-learning' and README: 'Python suite to construct benchmark machine learning datasets from the MIMIC-III clinical database... benchmark datasets cover ... classification, time series classification, regression, multilabel sequence classification.'