An LLM-based system that connects massive data with human requests:
it autonomously manages, processes, analyzes, predicts and visualizes data.
- 2024.05: Data-Copilot was presented as an Oral at the ICLR 2024 Workshop on Large Language Models for Agents.
- 2023.06: We released the paper, the code and an online demo.
Data-Copilot is an LLM-based system that helps you address data-related tasks. It connects data sources from different domains with diverse user needs, and can autonomously manage, process, analyze, predict and visualize data. When a request is received, it transforms raw data into the informative results that best match the user's intent.
- ⭐ Designer: Data-Copilot independently designs versatile interface tools with different functions through self-request and iterative refinement.
- ⭐ Dispatcher: Data-Copilot invokes the corresponding interfaces sequentially or in parallel, and transforms raw data from heterogeneous sources into graphics, tables and text, without human assistance.
This repository releases Data-Copilot for the Chinese financial market: stocks, funds, economic data, company financial data and live news.
Paper: Data-Copilot: Bridging Billions of Data and Humans with Autonomous Workflow
Wenqi Zhang, Yongliang Shen, Zeqi Tan, Guiyang Hou, Weiming Lu, Yueting Zhuang
Watch the demo video. Data-Copilot queries and predicts data autonomously:
Supported models and data sources:
| Model | CHN Stock | CHN Fund | CHN Economic data | CHN Financial data |
|---|---|---|---|---|
| OpenAI GPT-3.5 | ✓ | ✓ | ✓ | ✓ |
| Azure GPT-3.5 | ✓ | ✓ | ✓ | ✓ |
| Qwen-72B-Chat | ✓ | ✓ | ✓ | ✓ |
Data-Copilot was built on GPT-3.5, whose input was limited to 4k tokens at the time, so this release covers Chinese stocks, funds and economic data. Data from foreign financial markets may be supported in the future.
Python 3.8–3.10 is required (tested with Python 3.9 and 3.10).
git clone https://github.com/ZJU-OmniAI/Data-Copilot.git
cd Data-Copilot
conda create -n data-copilot python=3.10 -y
conda activate data-copilot
pip install -r requirements.txtData-Copilot reads its keys from environment variables:
| Variable | Needed for | Get it from |
|---|---|---|
TUSHARE_TOKEN |
all financial data | Tushare |
OPENAI_KEY |
GPT in main.py |
OpenAI |
DASHSCOPE_API_KEY |
Qwen in main.py |
Alibaba Cloud Bailian |
export TUSHARE_TOKEN="your-tushare-token"
export OPENAI_KEY="sk-..."Some Tushare interfaces used here are only open to accounts with enough credits (积分). In the web demo, you enter the OpenAI or Azure-OpenAI key on the page instead.
python main.pymain.py answers the example request at the end of the file. To ask something else, edit instruction. To switch the LLM, set model to "gpt" or "qwen-chat-72b".
The model versions are set in lab_gpt4_call.py (gpt-3.5-turbo) and lab_llms_call.py (qwen-72b-chat). If your provider no longer offers them, change the model names there.
python app.pyThen open http://127.0.0.1:7860. An online version of this demo is on Hugging Face Spaces.
- Enter your OpenAI key and click OK. For Azure-OpenAI, enter the key, the API base and the deployment name (engine) instead. A paid OpenAI plan is recommended: with the rate limits of a free plan, the demo is very slow.
- Type your request, or pick one from the example boxes.
- Click Start. The Solving Step box shows the intermediate workflow. The final answer appears as text (Summary and Result), a chart and a table.
Note
Stock names are matched against a local copy of the stock list saved on 2023-04-21 (data/). Stocks that were listed or renamed after that date are not found.
Data-Copilot
├── main.py # the workflow: intent detection → task planning → tool calls → visualization → summary
├── app.py # Gradio web demo
├── tool.py # interface tools: data acquisition, processing, prediction and visualization
├── lab_gpt4_call.py # OpenAI and Azure-OpenAI calls
├── lab_llms_call.py # Qwen (DashScope) and GLM calls
├── lab_llm_local_call.py # local InternLM (optional; needs torch, transformers and modelscope)
├── prompt_lib/ # prompts and in-context demonstrations for each stage and task
├── tool_lib/ # interface descriptions shown to the LLM, one file per task
├── create_tool/ # atomic Tushare API descriptions used for interface design
├── data/ # local lists: A-share stocks, funds, SW2021 industries
├── fonts/ # SimHei font for Chinese text in charts
├── output/ # tables saved by print_save_table
├── docs/flowchart.md # flowchart of main.py
└── assets/ # figures used in this README
tool.py and tool_lib/ contain the interface tools obtained in the first phase (interface design). prompt_lib/ contains the prompts and in-context demonstrations.
main.py handles a request in five stages:
- Intent detection: rewrites the request into a precise instruction with concrete dates and indicators.
- Task planning: splits the instruction into a data task (
stock_task,fund_taskoreconomic_task) and avisualization_task. - Tool calls: the LLM plans a workflow of interface calls for the data task (step by step, in parallel, or in a loop), and Data-Copilot runs it.
- Visualization: a second workflow draws charts or prints tables from the results.
- Summary: the LLM explains the plan, the tools it used and the result.
A flowchart of main.py is also available.
Daily net inflow and cumulative inflow of northbound capital this year
If you find this work useful, please cite:
@article{zhang2023data,
title={Data-Copilot: Bridging Billions of Data and Humans with Autonomous Workflow},
author={Zhang, Wenqi and Shen, Yongliang and Tan, Zeqi and Hou, Guiyang and Lu, Weiming and Zhuang, Yueting},
journal={arXiv preprint arXiv:2306.07209},
year={2023}
}If you have any questions, please open an issue or email [email protected].





