Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
A machine learning tool for fishing entities
| Date | Stars |
|---|---|
| 2026-07-31 | 268 |
| 2026-08-05 | 268 |
| 2026-08-06 | 268 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
[](http://www.apache.org/licenses/LICENSE-2.0.html) [](https://readthedocs.org/projects/nerd/?badge=latest) [](https://archive.softwareheritage.org/browse/origin/?origin_url=https://github.com/kermitt2/entity-fishing) [](http://cloud.science-miner.com/nerd) [](https://hub.docker.com/r/grobid/entity-fishing/ "Docker Pulls") # entity-fishing *entity-fishing* performs the following tasks for 15 different languages: * general entity recognition and disambiguation against Wikidata in a raw text or partially-annotated text segment,  * general entity recognition and disambiguation against Wikidata at document level, in particular for a PDF with layout positioning and structure-aware annotations,  * search query disambiguation (the _short text_ mode) - below disambiguation of the search query "concrete pump sensor" in the service test console,  * weighted term vector disambiguation (a term being a phrase),  * interactive disambiguation in text editing mode (experimental).  # Documentation [Presentation of entity-fishing at WikiDataCon 2017](https://grobid.s3.amazonaws.com/presentations/29-10-2017.pdf) for some design, implementation descriptions, and some evaluations. The documentation of *entity-fishing* is available [here](http://nerd.readthedocs.io). *entity-fishing* uses a query DSL (entity disambiguation specific query language) documented [here](https://nerd.readthedocs.io/en/latest/restAPI.html). # Demo For testing purposes, a public entity-fishing demo server (current version) is available at the following address: [https://cloud.science-miner.com/nerd](https://cloud.science-miner.com/nerd) The query DSL and Web services are documented [here](https://nerd.readthedocs.io/en/latest/restAPI.html). _Warning_: Some quota and query limitation apply to the demo server! Please be courteous and do not overload the demo server. A [docker image](https://nerd.readthedocs.io/en/latest/docker.html) is available to help you deploying your own server. # Benchmarks  Evaluations above correspond to the "overall unnormalized accuracy" scenario in [BLINK](https://github.com/facebookresearch/BLINK#benchmarking-blink) and are limited to Named Entities. entity-fishing performs at 0.765 F-score, as compared to 0.8027 for BLINK, a fine-tuned BERT architectures. *entity-fishing* surpasses BLINK for the dataset AQUAINT, 0.891 vs. 0.8588, and MSNBC, 0.867 vs. 0.8509, despite being considerably faster and lighter than BLINK (see below). See the [evaluation documentation](https://nerd.readthedocs.io/en/latest/evaluation.html) and [Presentation of entity-fishing at WikiDataCon 2017](https://grobid.s3.amazonaws.com/presentations/29-10-2017.pdf) for more details. *entity-fishing* has been designed to be particularly fast for a full scale Wikidata-based entity disambiguation tool (in particular not limited to Named Entities). On a single server, depending on the concurrency, it is possible to process from 1.000-1.500 tokens per seconds with concurrency 1 to 5.000 tokens per second for instance with 6 concurrent requests in English (and up to 2 times faster for other languages). # Some use cases Some example of *entity-fishing* usages: * A [spaCy wrapper](https://spacy.io/universe/project/spacyfishing) for
Excerpt of 8,480 characters
Read on GitHubWould you bet a product on this? Bounded 0–100 and slow moving.
matched fp:4d48d97fdf3796db, llm:Repo topics: entity-disambiguation, knowledge-engineering, machine-learning, wikidata, wikipedia; description and README: ML tool for entity recognition and disambiguation against Wikidata, supports NER, document-level and search-query disambiguation, multilingual.
matched fp:4d48d97fdf3796db, llm:Repo topics: entity-disambiguation, knowledge-engineering, machine-learning, wikidata, wikipedia; description and README: ML tool for entity recognition and disambiguation against Wikidata, supports NER, document-level and search-query disambiguation, multilingual.
matched fp:4d48d97fdf3796db, llm:Repo topics: entity-disambiguation, knowledge-engineering, machine-learning, wikidata, wikipedia; description and README: ML tool for entity recognition and disambiguation against Wikidata, supports NER, document-level and search-query disambiguation, multilingual.