Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
NucliaDB, The AI Search database for RAG
| Date | Stars |
|---|---|
| 2026-07-24 | 718 |
| 2026-07-25 | 719 |
| 2026-07-28 | 719 |
| 2026-07-30 | 719 |
| 2026-08-06 | 719 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
15.0
growth rate 0.00%/day
[](CODE_OF_CONDUCT.md) [](LICENCE.md)    [](https://codecov.io/gh/nuclia/nucliadb) <p align="center"> <img src="docs/assets/images/nuclia_db_positiu.svg" alt="Nuclia" height="100"> </p> <h3 align="center">The AI Search Database.</h3> <h4 align="center"> <a href="https://docs.nuclia.dev/docs/management/nucliadb/intro">Quickstart</a> | <a href="https://docs.nuclia.dev/docs/">Nuclia Docs</a> | <a href="https://nuclia-community.slack.com">Community</a> </h4> NucliaDB is a robust database that allows storing and searching on unstructured data. It is an out of the box hybrid search database, utilizing vector, full text and graph indexes. NucliaDB is written in Rust and Python. We designed it to index large datasets and provide multi-teanant support. When utilizing NucliaDB with Nuclia cloud, you are able to the power of an NLP database without the hassle of data extraction, enrichment and inference. We do all the hard work for you. # Features - Store text, files, vectors, labels and annotations - Perform text searches and given a word or set of words, return resources in our database that contain them. - Perform semantic searches with vectors. For example, given a set of vectors, return the closest matches in our database. With NLP, this allows us to look for similar sentences without being constrained by exact keywords. - Export your data in a format compatible with most NLP pipelines (HuggingFace datasets, pytorch, etc) - Store original data, extracting and data pulled from the Understanding API - Index fields, paragraphs, and semantic sentences on index storage - Cloud data and insight extraction with the Nuclia Understanding API™ - Cloud connection to train ML models with Nuclia Learning API™ - Role based security system with upstream proxy authentication validation - Resources with multiple fields and metadata - Text/HTML/Markdown plain fields support - Field types: text, file, link, conversation - Storage layer (PostgreSQL) - Blob support with S3-compatible API, GCS and Azure Blob Storage - Replication of index storage - Distributed search - Cloud-native ## Architecture <p align="center"> <img src="docs/assets/images/nucliadb-arch-overview.png" alt="Architecture" width="500px" style="background-color: #fff"> </p> ## Quickstart Trying NucliaDB is super easy! You can extend your knowledge with the following readings: - [Quick start!](https://docs.nuclia.dev/docs/management/nucliadb/intro) - Read about what Knowledge boxes are in [our basic concepts](https://docs.nuclia.dev/docs/management/nucliadb/basics) section - [Upload your data](https://docs.nuclia.dev/docs/ingestion/intro) # 💬 Community - Chat with us in [Slack][slack] - 📝 [Blog Posts][blogs] - Follow us on [X][X] - Do you want to [work with us][linkedin]? # 🙋 FAQ ## How is NucliaDB different from traditional search engines like Elasticsearch or Solr? The core difference and advantage of NucliaDB is its architecture built from the ground up for unstructured data. Its vector index, keyword, graph and fuzzy search provide an API to use all extracted and extracted information from Nuclia, Understanding API and provides powerful NLP abilities to any application with low code and peace of mind. ## What license does NucliaDB use? NucliaDB is open-source under the GNU Affero General Public License Version 3 - AGPLv3. Fundamentally, this means that you are free to use NucliaDB for your project, as long as you don't modify NucliaDB. If you do, you have to make the m
Excerpt of 5,327 characters
Read on GitHub938
567
566
518
469
259
92
80
67
57
49
47
40
32
31
30
22
21
15
15
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:7b7d819068e21ebd, topic:vector-search, readme:vector index, readme:hybrid search
matched fp:7b7d819068e21ebd, topic:mlops
matched fp:7b7d819068e21ebd, topic:language-model
matched fp:7b7d819068e21ebd, topic:text-classification