AItonomy Foundation · 2026-09-16 · notable
ScienceIDE — scientific codebases become training grounds for agents
ScienceIDE turns real scientific code repositories into executable environments where AI agents are trained and graded on numerical correctness. The AItonomy Foundation also released three open PhAI-IDE models at 4B, 9B and 72B.
Real scientific code repositories, turned into environments where agents are graded on numerical correctness.
What is it?
ScienceIDE converts scientific code repositories — astrophysics, ocean modelling, neuroscience simulation — into programmable environments for AI agents. The AItonomy Foundation published 15 of its 64 environments plus the RL code and 30 of 85 ScienceIDE-Hard tasks, and released three open PhAI-IDE models at 4B, 9B and 72B trained on verified trajectories.
How does it work?
Agents repair injected defects in real scientific simulations by editing source code, and the reward comes from numerical correctness checks rather than matching a reference diff. Expert-defined scientific cases and acceptance criteria guide how each repository becomes a task, and the resulting verified trajectories feed supervised fine-tuning, reinforcement learning and evaluation from one shared foundation.
Why does it matter?
Decades of domain knowledge sit in scientific software that generic code benchmarks never touch, which the authors call the scientific experience bottleneck. The project reports the PhAI-IDE family also improving on general tests — 6.99 points on CodeXGLUE defect detection and 36.00 points on BBH Word Sorting — which it reads as evidence that scientific experience transfers to broader coding and reasoning.
Who is it for?
scientific-agent and RL researchers
Try it
https://huggingface.co/collections/AItonomy/scienceide-model-series