█

AI/TLDR

Roboflow Maestro

One CLI and one Python API for fine-tuning vision-language models — Florence-2, PaliGemma 2 and Qwen2.5-VL — over a single JSONL dataset format, with LoRA, QLoRA and graph freezing

Fine-Tuning FrameworksOpen source
Latest
1.0.0
Updated
5 Feb 2025
Language
Python
License
Apache-2.0
$pip install "maestro[paligemma_2]"

What's new

1.0.05 Feb 2025

Added support for Florence-2, PaliGemma 2 and Qwen2.5-VL together with LoRA, QLoRA and graph freezing to keep hardware requirements in check, and consolidated training behind a single CLI and SDK over one consistent JSONL data format.

Overview

Maestro is Roboflow's fine-tuning tool for multimodal models. Each vision-language model arrives with its own training script, its own expected data layout and its own set of half-documented hyperparameters, and most of the work in adapting one is reconciling those differences rather than training anything. Maestro absorbs that: it handles configuration, data loading, reproducibility and the training-loop setup, and exposes ready-to-use recipes for Florence-2, PaliGemma 2 and Qwen2.5-VL behind a shared interface.

There are two ways in, and they run the same core module. The CLI takes the parameters you would actually change — dataset location, epochs, batch size, optimisation strategy, metrics — as flags, which makes a run easy to script and easy to diff. The Python API takes the same values as a config dictionary for when you want more control over the surrounding code. Either way the data goes in as one consistent JSONL format across every supported model, so switching which model you are fine-tuning does not mean reshaping the dataset.

Hardware requirements are managed through the optimisation strategy rather than by asking for a bigger GPU. The 1.0.0 release added LoRA, QLoRA and graph freezing across the supported models, and the project publishes Colab cookbooks that fine-tune on free hardware: PaliGemma 2 (3B) and Qwen2.5-VL (3B) for JSON data extraction with LoRA and QLoRA respectively, plus object-detection recipes for Florence-2 (0.9B) and Qwen2.5-VL (7B) that it marks as experimental. Because some models have clashing requirements, dependencies are installed per model as extras and the project recommends a dedicated Python environment for each.

What it does

  • Ready-made fine-tuning recipes for Florence-2, PaliGemma 2 and Qwen2.5-VL behind one interface
  • A single CLI and a matching Python API over the same core modules, so a scripted run and an embedded one behave identically
  • One consistent JSONL dataset format across every supported model
  • LoRA, QLoRA and graph freezing to keep fine-tuning within reach of modest hardware
  • Configuration, data loading, reproducibility and training-loop setup handled by the library rather than copied between scripts
  • Colab cookbooks for object detection and JSON data extraction that run on free hardware

Getting started

Dependencies are installed per model as extras, because some of the supported models have clashing requirements — the project recommends a dedicated Python environment for each one you train.

Install the extras for the model you are fine-tuning

bashbash
pip install "maestro[paligemma_2]"

Train from the command line

The subcommand names the model; the flags are the parameters worth changing. Optimisation strategy is where hardware cost is controlled — qlora is the cheapest of the three.

bashbash
maestro paligemma_2 train \
  --dataset "dataset/location" \
  --epochs 10 \
  --batch-size 4 \
  --optimization_strategy "qlora" \
  --metrics "edit_distance"

Or drive the same run from Python

Import train from the corresponding model's core module and pass the same values as a dictionary. The core module takes care of reproducibility, data preparation and training setup.

pythonpython
from maestro.trainer.models.paligemma_2.core import train

config = {
    "dataset": "dataset/location",
    "epochs": 10,
    "batch_size": 4,
    "optimization_strategy": "qlora",
    "metrics": ["edit_distance"]
}

train(config)

Commands and code are distilled from the project's own documentation — always check the official repo for the latest.

When to use it

  • Reach for it when a general-purpose vision-language model nearly does your task and needs adapting to your own images and labels
  • Reach for it when you want to extract structured JSON from documents or screenshots and would rather fine-tune a small VLM than prompt a large one
  • Reach for it when you want to compare two or three VLM backbones on the same dataset without rewriting the data pipeline for each
  • Reach for it when the available hardware is one modest GPU and the run has to fit inside it via LoRA, QLoRA or graph freezing

How Roboflow Maestro compares

Roboflow Maestro alongside other open-source fine-tuning frameworks tools AI/TLDR tracks, ranked by GitHub stars.

ToolStarsWhat it does
Unsloth★ 76.9kA library that speeds up LoRA and QLoRA fine-tuning while cutting memory use, aimed at training models on a single GPU.
LLaMA-Factory★ 75.1kAn end-to-end training suite with a web UI that covers pre-training, supervised fine-tuning, and RLHF for hundreds of LLMs and multimodal models.
PEFT★ 21.7kHugging Face's library of parameter-efficient fine-tuning methods such as LoRA, DoRA, and prompt tuning that train small adapters instead of full models.
FinGPT★ 21.3kFinGPT is an open-source project of financial LLMs, fine-tuned with LoRA on news and tweet data for tasks like sentiment analysis, relation extraction, and stock-move forecasting.
ms-swift★ 15.7kModelScope's framework for fine-tuning and deploying 600+ LLMs and 300+ multimodal models, supporting PEFT and full-parameter SFT, DPO, and GRPO.
LitGPT★ 13.7kAn open-source toolkit from Lightning AI to pretrain, finetune, and serve 20+ large language models, each written from scratch for speed and full control.
Axolotl★ 12.5kA config-driven tool for fine-tuning and post-training open LLMs that supports SFT, LoRA/QLoRA, DPO, GRPO, and multi-GPU training across many model families.
Roboflow Maestro★ 2.7kOne CLI and one Python API for fine-tuning vision-language models — Florence-2, PaliGemma 2 and Qwen2.5-VL — over a single JSONL dataset format, with LoRA, QLoRA and graph freezing