VisionPack CLI Guide
VisionPack is a CLI-first DatasetOps tool for computer vision datasets. It supports local projects across classification, detection, instance segmentation, and keypoints; YOLO / COCO / ImageFolder import and export; declarative multi-source sync; validation (including near-duplicate and cross-split leakage detection); statistics; deterministic splits; snapshots and snapshot diffs; and archive and WebDataset training packs.
For architecture and the roadmap see ARCHITECTURE.md.
Install
pip install visionpack
Cloud backends are optional extras: pip install "visionpack[s3]" (also [gcs],
[azure]). See Installation for details.
For development from source with uv:
uv sync
uv run vp --help
uv run python -m visionpack --help # or run the module directly
The examples below use the bare vp command; prefix with uv run when working
from a source checkout.
Every pipeline command also takes --json and prints one machine-readable,
schema-versioned document to stdout — the supported way to drive VisionPack
from another program. See JSON Output.
Initialize A Dataset
Create a VisionPack project in the current directory:
vp init --name factory-defects --task detection
This creates a git-like layout — just the manifest and a control directory:
visionpack.yaml
.vp/
db/ # local index (index.db, SQLite)
objects/ # content-addressed assets (sha256)
snapshots/ # versioned snapshots
The visionpack.yaml file is the declarative dataset manifest and, together with
the .vp/ index, is the single source of truth. Output directories such as
exports/ and reports/ are created on demand by the commands that write them, so
the project root stays clean.
Import A YOLO Dataset
Supported YOLO layouts include image and label files side by side:
raw/
classes.txt
img001.jpg
img001.txt
And split-style image/label folders:
raw/
classes.txt
images/
img001.jpg
labels/
img001.txt
Import the dataset:
vp import ./raw # format auto-detected from the layout
vp import ./raw --format yolo # or say it explicitly
--format defaults to auto: a .json annotations file (or a directory with
an instances-style JSON at its root) is COCO, .txt labels or
classes.txt/data.yaml mean YOLO, and images living only under
folder-per-class subdirectories mean ImageFolder. When the layout is ambiguous
the command asks you to pass --format instead of guessing.
By default, VisionPack ingests images into .vp/objects/sha256 and indexes them by content hash. You can choose another copy mode:
vp import ./raw --format yolo --copy hardlink
vp import ./raw --format yolo --copy reference
Available copy modes are copy, move, hardlink, reference, and ingest.
Each successful import is also recorded as a source in visionpack.yaml, so the
manifest reflects where your data came from and you can re-pull it later with
vp sync. Pass --no-record to skip this (for a one-off/throwaway import), or
--name to control the recorded source name.
Multi-source sync
Declare images and labels — even when they live in different folders or repos — in a
sources: block and reconcile them with vp sync:
sources:
- name: camera-A
format: yolo
images: ./repoA/images
labels: ./repoB
match: stem # pair by filename; use `relpath` for parallel trees
copy: ingest
vp sync --dry-run # preview found / matched / unmatched / classes per source
vp sync # ingest; idempotent, records per-asset provenance
vp sync --source camera-A # sync just one source
vp sync --jobs 32 # concurrent transfers per source (remote defaults to 16+)
Sources can also live in object stores. Remote URIs go anywhere a local path
would, and a target: lets copy mode land objects server-side in a
content-addressed bucket without downloading them:
target: s3://my-bucket/datasets/factory-defects
sources:
- name: camera-A
format: yolo
images: s3://my-bucket/raw/camera-a/images
labels: s3://my-bucket/raw/camera-a/labels
copy: copy
See Cloud Sync for credentials, copy modes, and streaming export.
Validate
Run validation:
vp validate
Strict mode treats missing annotations as errors:
vp validate --strict
Write a JSON validation report:
vp validate --report reports/validation.json
The current validator checks image readability, missing annotations, orphan labels, unknown classes, invalid boxes, boxes outside image bounds, duplicate exact assets, and split leakage.
Audit Label Health (vp audit)
vp validate catches labels that are invalid; vp audit finds labels that
are valid but suspicious — the ones that usually turn out to be annotation
mistakes:
vp audit
It reports duplicate boxes (the same object labeled twice), degenerate (tiny) boxes, boxes pinned to two or more image borders, whole-image boxes, extreme aspect-ratio outliers, rare classes, and dataset-level class imbalance.
Findings are advisory (exit code 0). Gate CI on them explicitly:
vp audit --fail-on-findings
Thresholds can be tuned per run (--min-box-px, --duplicate-iou,
--max-aspect-ratio, --imbalance-ratio, --min-class-count) or persisted in
visionpack.yaml:
validation:
audit:
min_box_px: 12
duplicate_iou: 0.85
Like every pipeline command, vp audit --json prints a machine-readable
envelope with per-code counts and the full findings list.
Show Statistics
Print a summary:
vp stats
Class distribution:
vp stats --by class
JSON output:
vp stats --json
Create Snapshots
Create a reproducible dataset snapshot:
vp snapshot create -m "initial import"
List snapshots:
vp snapshot list
Show one snapshot:
vp snapshot show v1
Snapshots store hashes for the manifest, assets, annotations, splits, and summary stats. They are written to .vp/snapshots/.
Lineage tags (which dataset trained this model?)
After a training run, stamp the snapshot it consumed:
vp snapshot tag v4 trained:run-812
Tags are free-form (the key:value convention is just a convention), show up
in vp snapshot list/show, and can be removed with --remove. From the SDK,
ds.snapshots_by_tag("trained:") lists every version any run trained on — so
“which dataset produced this model?” is a lookup, not archaeology.
Diff Snapshots
Compare two snapshots:
vp diff v1 v2
JSON diff:
vp diff v1 v2 --json
The diff reports added and removed assets, added/removed/modified annotations, class changes, split changes, and before/after stats.
Distribution drift
--drift adds a class-distribution comparison: per-class object counts and
distribution-share deltas (biggest movers first), plus KL and Jensen–Shannon
divergence as single drift scores a CI job can threshold:
vp diff v1 v2 --drift
vp diff v1 v2 --drift --json # adds a "drift" object to the diff payload
Because it derives from the stats frozen inside each snapshot, the drift between two versions is reproducible forever.
Export YOLO
Export the indexed dataset back to YOLO format:
vp export --format yolo --output exports/yolo-v1
The export writes:
exports/yolo-v1/
images/
labels/
classes.txt
data.yaml
Exports never duplicate bytes: local images are hardlinked from the
content-addressed store, and cloud-backed images are written to a
manifest.jsonl (image → object URI) for streaming instead of being downloaded.
See Cloud Sync.
Segmentation projects export YOLO-seg polygon labels automatically (force either
way with --seg / --no-seg), and --format masks writes semantic masks —
8-bit PNGs whose pixel value is the class index (0 = background, mapping recorded
in classes.txt):
vp export --format yolo --output exports/seg-v1 --split # YOLO-seg labels
vp export --format masks --output exports/masks-v1 --split # class-index PNGs
Evaluate A Model (vp eval)
Score model predictions against the labels of a split’s set (the test set by default) — a locked split plus a snapshot makes the number a reproducible benchmark:
vp eval runs/predict/labels --format yolo # Ultralytics save_txt output
vp eval predictions.json # vp-native or COCO JSON
vp eval predictions.json --split default --set val --json
Detection, segmentation, and keypoints report per-class AP@50, mAP@50,
mAP@50-95, and precision/recall at a confidence threshold (--conf);
classification reports accuracy, per-class precision/recall/F1, and a confusion
matrix. Predictions reference images by asset id (what exports name files) or by
original filename; unresolvable references are reported, never dropped silently.
Autolabel & The Annotation Queue
vp autolabel persists confident predictions as annotations. Model labels are
recorded with source.type = "model", so they stay distinguishable from human
labels. Only unlabeled assets are touched unless you pass --replace:
vp autolabel predictions.json --min-confidence 0.6
vp queue ranks images by how much a human label would help (active learning):
unlabeled images first — ordered by model uncertainty when predictions are given
— and with --include-labeled it audits existing labels for ground-truth /
prediction disagreement (possible missing or stale labels):
vp queue --predictions predictions.json --limit 20
vp queue --predictions predictions.json --include-labeled --json
Together these close the model-in-the-loop cycle: export → train/predict →
vp eval (measure) → vp autolabel (label the easy images) → vp queue
(send the hard ones to humans) → re-train.
Pack Archive
Create a compressed archive package:
vp pack --profile archive
Or choose an output path:
vp pack --profile archive --output exports/archive/factory-defects.tar.zst
The archive includes:
visionpack.yaml.vp/db/index.db.vp/snapshots/*.json- content-addressed assets
pack.jsonwith pack metadata and dataset stats
Python API
The supported programmatic surface is the SDK — the whole CLI workflow behind one Python class, with the same locking and result shapes (see the Python SDK page):
from visionpack.sdk import VisionPackClient
ds = VisionPackClient.open(".")
print(ds.name, len(ds))
ds.validate()
ds.export("./exports/yolo", format="yolo", split="default")
The lower-level visionpack.Dataset / Project handle stays available as
ds.project for anything the facade doesn’t cover yet.
Current Limitations
- segmentation metrics in
vp evaluse each polygon’s enclosing box (mask IoU is planned); YOLO-pose import/export and a dedicated keypoint importer are not implemented yet vp annotateis scaffolded but not implemented yet- cloud sync (S3/GCS/Azure) is same-provider in v1 — cross-cloud transfer
(S3↔GCS) and remote COCO/ImageFolder sync are planned;
packis local-only - the local index is SQLite (
index.db);statsand exports stream records (flat RAM at scale), whilevalidate,split, dedup, and the WebDataset pack still load the full set (next streaming pass — see ARCHITECTURE.md)