functime-org/functime
quality grade C, 62 out of 100Time-series machine learning at scale. Built with Polars for embarrassingly parallel feature extraction and forecasts on panel data.
- stars
- 1.2k
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
ETL, document parsing, web crawling, synthetic data generation, annotation and dataset curation tooling.
Signals: data-engineering, etl, data-pipeline, web-scraping, crawler, data-labeling, annotation-tool, synthetic-data
473 results
Time-series machine learning at scale. Built with Polars for embarrassingly parallel feature extraction and forecasts on panel data.
A library that uses hardware acceleration to load sequences of video frames to facilitate machine learning training
Python Fast Dataflow programming framework for Data pipeline work( Web Crawler,Machine Learning,Quantitative Trading.etc)
C#/F# bindings for NumPy - a fundamental library for scientific computing, machine learning and AI
Python Client and Toolkit for DataFrames, Big Data, Machine Learning and ETL in Elasticsearch
SFrame: Scalable tabular and graph data-structures built for out-of-core data analysis and machine learning.
Code repo for the book "Feature Engineering for Machine Learning," by Alice Zheng and Amanda Casari, O'Reilly 2018
DataFrames for Go: For statistics, machine-learning, and data manipulation/exploration
Well-documented Python demonstrations for spatial data analytics, geostatistical and machine learning to support my courses.
Apache Spark & Python (pySpark) tutorials for Big Data Analysis and Machine Learning as IPython / Jupyter notebooks
Library for exploring and validating machine learning data
MLDB is the Machine Learning Database
Machine Learning Time-Series Platform
Machine learning for transportation data imputation and prediction.
A Data Engineering & Machine Learning Knowledge Hub
A data pipeline framework for machine learning
Cleanlab's open-source library is the standard data-centric AI package for data quality and machine learning with messy, real-world data and labels.
A Python Package to Tackle the Curse of Imbalanced Datasets in Machine Learning
汉字字符特征提取器 (featurizer),提取汉字的特征(发音特征、字形特征)用做深度学习的特征 | A Chinese character feature extractor, which extracts the features of Chinese characters (pronunciation features, glyph features) as features for deep learning
This is a reading list for deep learning for OCR
An efficient video loader for deep learning with smart shuffling that's super easy to digest
Generate text line images for training deep learning OCR models
Deep Learning Pipelines for Apache Spark
This library augments road images to introduce various real world scenarios that pose challenges for training neural networks of Autonomous vehicles. Automold is created to train CNNs in specific weather and road conditions.
24,523 repositories in the index in total.