Overview
Maestro is Roboflow's fine-tuning tool for multimodal models. Each vision-language model arrives with its own training script, its own expected data layout and its own set of half-documented hyperparameters, and most of the work in adapting one is reconciling those differences rather than training anything. Maestro absorbs that: it handles configuration, data loading, reproducibility and the training-loop setup, and exposes ready-to-use recipes for Florence-2, PaliGemma 2 and Qwen2.5-VL behind a shared interface.
There are two ways in, and they run the same core module. The CLI takes the parameters you would actually change — dataset location, epochs, batch size, optimisation strategy, metrics — as flags, which makes a run easy to script and easy to diff. The Python API takes the same values as a config dictionary for when you want more control over the surrounding code. Either way the data goes in as one consistent JSONL format across every supported model, so switching which model you are fine-tuning does not mean reshaping the dataset.
Hardware requirements are managed through the optimisation strategy rather than by asking for a bigger GPU. The 1.0.0 release added LoRA, QLoRA and graph freezing across the supported models, and the project publishes Colab cookbooks that fine-tune on free hardware: PaliGemma 2 (3B) and Qwen2.5-VL (3B) for JSON data extraction with LoRA and QLoRA respectively, plus object-detection recipes for Florence-2 (0.9B) and Qwen2.5-VL (7B) that it marks as experimental. Because some models have clashing requirements, dependencies are installed per model as extras and the project recommends a dedicated Python environment for each.
What it does
- Ready-made fine-tuning recipes for Florence-2, PaliGemma 2 and Qwen2.5-VL behind one interface
- A single CLI and a matching Python API over the same core modules, so a scripted run and an embedded one behave identically
- One consistent JSONL dataset format across every supported model
- LoRA, QLoRA and graph freezing to keep fine-tuning within reach of modest hardware
- Configuration, data loading, reproducibility and training-loop setup handled by the library rather than copied between scripts
- Colab cookbooks for object detection and JSON data extraction that run on free hardware
Getting started
Dependencies are installed per model as extras, because some of the supported models have clashing requirements — the project recommends a dedicated Python environment for each one you train.
Install the extras for the model you are fine-tuning
pip install "maestro[paligemma_2]"Train from the command line
The subcommand names the model; the flags are the parameters worth changing. Optimisation strategy is where hardware cost is controlled — qlora is the cheapest of the three.
maestro paligemma_2 train \
--dataset "dataset/location" \
--epochs 10 \
--batch-size 4 \
--optimization_strategy "qlora" \
--metrics "edit_distance"Or drive the same run from Python
Import train from the corresponding model's core module and pass the same values as a dictionary. The core module takes care of reproducibility, data preparation and training setup.
from maestro.trainer.models.paligemma_2.core import train
config = {
"dataset": "dataset/location",
"epochs": 10,
"batch_size": 4,
"optimization_strategy": "qlora",
"metrics": ["edit_distance"]
}
train(config)Commands and code are distilled from the project's own documentation — always check the official repo for the latest.
When to use it
- Reach for it when a general-purpose vision-language model nearly does your task and needs adapting to your own images and labels
- Reach for it when you want to extract structured JSON from documents or screenshots and would rather fine-tune a small VLM than prompt a large one
- Reach for it when you want to compare two or three VLM backbones on the same dataset without rewriting the data pipeline for each
- Reach for it when the available hardware is one modest GPU and the run has to fit inside it via LoRA, QLoRA or graph freezing
How Roboflow Maestro compares
Roboflow Maestro alongside other open-source fine-tuning frameworks tools AI/TLDR tracks, ranked by GitHub stars.
| Tool | Stars | What it does |
|---|---|---|
| Unsloth | ★ 76.9k | A library that speeds up LoRA and QLoRA fine-tuning while cutting memory use, aimed at training models on a single GPU. |
| LLaMA-Factory | ★ 75.1k | An end-to-end training suite with a web UI that covers pre-training, supervised fine-tuning, and RLHF for hundreds of LLMs and multimodal models. |
| PEFT | ★ 21.7k | Hugging Face's library of parameter-efficient fine-tuning methods such as LoRA, DoRA, and prompt tuning that train small adapters instead of full models. |
| FinGPT | ★ 21.3k | FinGPT is an open-source project of financial LLMs, fine-tuned with LoRA on news and tweet data for tasks like sentiment analysis, relation extraction, and stock-move forecasting. |
| ms-swift | ★ 15.7k | ModelScope's framework for fine-tuning and deploying 600+ LLMs and 300+ multimodal models, supporting PEFT and full-parameter SFT, DPO, and GRPO. |
| LitGPT | ★ 13.7k | An open-source toolkit from Lightning AI to pretrain, finetune, and serve 20+ large language models, each written from scratch for speed and full control. |
| Axolotl | ★ 12.5k | A config-driven tool for fine-tuning and post-training open LLMs that supports SFT, LoRA/QLoRA, DPO, GRPO, and multi-GPU training across many model families. |
| Roboflow Maestro | ★ 2.7k | One CLI and one Python API for fine-tuning vision-language models — Florence-2, PaliGemma 2 and Qwen2.5-VL — over a single JSONL dataset format, with LoRA, QLoRA and graph freezing |