Overview
DeepAnalyze, from the Renmin University of China data lab with Tsinghua collaborators, is an agentic LLM built specifically for data work rather than a general assistant pointed at a CSV. Given a task and a set of files, it runs the whole pipeline without a human in the loop: data preparation, analysis, modelling, visualisation and report generation. It also handles open-ended data research — explore several sources and produce an analyst-grade written report rather than answer one question.
It is explicitly source-agnostic. Structured inputs (databases, CSV, Excel), semi-structured inputs (JSON, XML, YAML) and unstructured inputs (TXT, Markdown) can all be listed in one prompt, and the agent discovers what is in them before deciding how to use them. The whole stack is open: the DeepAnalyze-8B checkpoint on Hugging Face and ModelScope, the training data as DataScience-Instruct-500K, the inference and training code, and a demo — so you can deploy it as-is or continue training your own variant.
Four interfaces ship with the project: a WebUI, a WebUI v2 built on the DA-Studio system (accepted to the VLDB 2026 demonstration track) with Docker-sandboxed code execution, a JupyterUI, and a CLI. Around the core model the group has published companion work — DeepPrep for autonomous data preparation, CoDA-Bench for evaluating code agents on data-intensive tasks, and SkillAdam for automatically improving agent skills — each in its own repository.
What it does
- Runs the full data-science pipeline end to end — preparation, analysis, modelling, visualisation and report writing
- Open-ended data research across many sources, producing analyst-grade written reports rather than single answers
- Handles structured (databases, CSV, Excel), semi-structured (JSON, XML, YAML) and unstructured (TXT, Markdown) inputs in one task
- Fully open release: DeepAnalyze-8B weights, DataScience-Instruct-500K training data, inference and training code
- Four front ends — WebUI, WebUI v2 (DA-Studio) with Docker-sandboxed execution, JupyterUI and a CLI
- Served through vLLM behind an OpenAI-compatible endpoint, with 4-bit and 8-bit quantised variants for 16–24 GB GPUs
- Curriculum-based agentic training recipes built on ms-swift and SkyRL for training your own variant
Getting started
DeepAnalyze runs the DeepAnalyze-8B checkpoint locally through vLLM, or you can request an API key from the project instead of hosting it. The README recommends keeping the inference and training environments separate to avoid dependency conflicts.
Set up the environment
Python 3.12 with torch, transformers and vllm 0.8.5 or newer. requirements.txt lists the minimal inference dependencies.
conda create -n deepanalyze python=3.12 -y
conda activate deepanalyze
pip install -r requirements.txtGet the model
Download DeepAnalyze-8B from Hugging Face (RUC-DataLab/DeepAnalyze-8B) or ModelScope. Use ./quantize.py to produce a 4-bit or 8-bit copy if your GPU has under 24 GB.
Serve it with vLLM
Pick max-model-len from the memory table in the README — for example 49152 with FP8 KV cache on a 16 GB card, or 131072 on 40 GB and up. The service answers at http://localhost:8000/v1/completions.
python -m vllm.entrypoints.openai.api_server \
--model /path/to/deepanalyze/4bit \
--served-model-name DeepAnalyze-8B \
--max-model-len 49152 \
--gpu-memory-utilization 0.95 \
--port 8000 \
--kv-cache-dtype fp8 \
--trust-remote-codeRun a task
Describe the task and list the data sources; the agent explores them itself. Any number and any mix of file types is allowed.
from deepanalyze import DeepAnalyzeVLLM
prompt = """# Instruction
Generate a data science report.
# Data
File 1: {"name": "person.csv", "size": "10.6KB"}
File 2: {"name": "enlist.csv", "size": "6.7KB"}
File 3: {"name": "male.xlsx", "size": "8.8KB"}
"""Train your own variant
Training uses a separate environment. Install the vendored ms-swift and SkyRL trees in editable mode, then follow the curriculum-based agentic training scripts.
(cd ./deepanalyze/ms-swift/ && pip install -e .)
(cd ./deepanalyze/SkyRL/ && pip install -e .)Commands and code are distilled from the project's own documentation — always check the official repo for the latest.
When to use it
- Reach for it when a folder of mixed CSVs, spreadsheets and JSON needs exploring and you do not know yet what the question is
- Reach for it to generate a repeatable analyst-style report from a fixed set of sources rather than a one-off chat answer
- Reach for it when the data must stay on your own hardware — the weights, code and training set are all open
- Reach for it as a base model to fine-tune on your organisation's own data-science workflows
How DeepAnalyze compares
DeepAnalyze alongside other open-source agent frameworks & builders tools AI/TLDR tracks, ranked by GitHub stars.
| Tool | Stars | What it does |
|---|---|---|
| DeepSeek Harness | ★ 243k | DeepSeek AI's open-source agent harness (dsh), built on Cordis, where models, tools, skills, sessions, sandboxes, storage and the UI are all plugins composed through profiles. |
| AutoGPT | ★ 188k | One of the earliest autonomous agent projects, now a platform for building and running agents from reusable blocks and workflows. |
| DeerFlow | ★ 83.4k | ByteDance's open-source super agent harness built on LangGraph: skills, sub-agents, sandboxes, a filesystem and long-term memory for long-horizon research, coding and content tasks. |
| nanobot | ★ 48.8k | Lightweight self-hosted personal AI agent framework in Python, with a WebUI, terminal and chat-app channels, tools, long-term memory, MCP and scheduled automations. |
| LangGraph | ★ 42.7k | A library from the LangChain team for building stateful, graph-based agent workflows with explicit control over steps, memory, and human-in-the-loop checkpoints. |
| Agno | ★ 42.5k | A fast Python framework (formerly Phidata) for building agents with memory, tools, and multimodal inputs, plus a runtime for deploying them in production. |
| AgentGPT | ★ 36.3k | AgentGPT lets you name a custom AI, give it a goal, and watch it plan tasks, run them, and learn from the results, all from a web browser. |
| DeepAnalyze | ★ 4.7k | Agentic LLM that runs a whole data-science pipeline over your files and databases |