wkunzhi/Python3-Spider
quality grade F, 32 out of 100Python爬虫实战 - 模拟登陆各大网站 包含但不限于:滑块验证、拼多多、美团、百度、bilibili、大众点评、淘宝,如果喜欢请start ❤️
- stars
- 3.4k
- stars gained this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
ETL, document parsing, web crawling, synthetic data generation, annotation and dataset curation tooling.
Signals: data-engineering, etl, data-pipeline, web-scraping, crawler, data-labeling, annotation-tool, synthetic-data
473 results
Python爬虫实战 - 模拟登陆各大网站 包含但不限于:滑块验证、拼多多、美团、百度、bilibili、大众点评、淘宝,如果喜欢请start ❤️
一些非常有趣的python爬虫例子,对新手比较友好,主要爬取淘宝、天猫、微信、微信读书、豆瓣、QQ等网站。(Some interesting examples of python crawlers that are friendly to beginners. )
Flock: A Low-Cost Streaming Query Engine on FaaS Platforms
A serverless architecture for orchestrating ETL jobs in arbitrarily-complex workflows using AWS Step Functions and AWS Lambda.
Enterprise-grade, production-hardened, serverless data lake on AWS
All-in-One Development Tool based on PaddlePaddle
A curated list of awesome synthetic data tools (open source and commercial).
Generative adversarial training for generating synthetic tabular data.
Repository for Paper: Cross-Domain Complementary Learning Using Pose for Multi-Person Part Segmentation (TCSVT20)
Code used to generate synthetic scenes and bounding box annotations for object detection. This was used to generate data used in the Cut, Paste and Learn paper
OpenXRLab Synthetic Data Rendering Toolbox
Random dataframe and database table generator
Official project website for the CVPR 2020 paper (Oral Presentation) "Cascaded deep monocular 3D human pose estimation wth evolutionary training data"
Synthetic Minority Over-Sampling Technique for Regression
Genalog is an open source, cross-platform python package allowing generation of synthetic document images with custom degradations and text alignment capabilities.
A novel approach for synthesizing tabular data using pretrained large language models
plait.py - a fake data modeler
Generate relevant synthetic data quickly for your projects. The Databricks Labs synthetic data generator (aka `dbldatagen`) may be used to generate large simulated / synthetic data sets for test, POCs, and other uses in Databricks environments including in Delta Live Tables pipelines
A tool that uses advanced Monte Carlo simulations and Turbit parallel processing to create possible Bitcoin prediction scenarios.
Draw a store, generate LLM personas, and watch them shop — an isometric 3D sandbox for synthetic-consumer experiments.
A curated list of awesome projects which use Machine Learning to generate synthetic content.
A library to model multivariate data using copulas.
A library for generating and evaluating synthetic tabular data for privacy, fairness and data augmentation.
Configurable Generation of Synthetic Schemas and Knowledge Graphs at Your Fingertips
24,523 repositories in the index in total.