VisionPack CLI Guide

VisionPack is a CLI-first DatasetOps tool for computer vision datasets. It supports local projects across classification, detection, instance segmentation, and keypoints; YOLO / COCO / ImageFolder import and export; declarative multi-source sync; validation (including near-duplicate and cross-split leakage detection); statistics; deterministic splits; snapshots and snapshot diffs; and archive and WebDataset training packs.

For architecture and the roadmap see ARCHITECTURE.md.

Install

pip install visionpack

Cloud backends are optional extras: pip install "visionpack[s3]" (also [gcs], [azure]). See Installation for details.

For development from source with uv:

uv sync
uv run vp --help
uv run python -m visionpack --help   # or run the module directly

The examples below use the bare vp command; prefix with uv run when working from a source checkout.

Every pipeline command also takes --json and prints one machine-readable, schema-versioned document to stdout — the supported way to drive VisionPack from another program. See JSON Output.

Initialize A Dataset

Create a VisionPack project in the current directory:

vp init --name factory-defects --task detection

This creates a git-like layout — just the manifest and a control directory:

visionpack.yaml
.vp/
  db/          # local index (index.db, SQLite)
  objects/     # content-addressed assets (sha256)
  snapshots/   # versioned snapshots

The visionpack.yaml file is the declarative dataset manifest and, together with the .vp/ index, is the single source of truth. Output directories such as exports/ and reports/ are created on demand by the commands that write them, so the project root stays clean.

Import A YOLO Dataset

Supported YOLO layouts include image and label files side by side:

raw/
  classes.txt
  img001.jpg
  img001.txt

And split-style image/label folders:

raw/
  classes.txt
  images/
    img001.jpg
  labels/
    img001.txt

Import the dataset:

vp import ./raw                 # format auto-detected from the layout
vp import ./raw --format yolo   # or say it explicitly

--format defaults to auto: a .json annotations file (or a directory with an instances-style JSON at its root) is COCO, .txt labels or classes.txt/data.yaml mean YOLO, and images living only under folder-per-class subdirectories mean ImageFolder. When the layout is ambiguous the command asks you to pass --format instead of guessing.

By default, VisionPack ingests images into .vp/objects/sha256 and indexes them by content hash. You can choose another copy mode:

vp import ./raw --format yolo --copy hardlink
vp import ./raw --format yolo --copy reference

Available copy modes are copy, move, hardlink, reference, and ingest.

Each successful import is also recorded as a source in visionpack.yaml, so the manifest reflects where your data came from and you can re-pull it later with vp sync. Pass --no-record to skip this (for a one-off/throwaway import), or --name to control the recorded source name.

Multi-source sync

Declare images and labels — even when they live in different folders or repos — in a sources: block and reconcile them with vp sync:

sources:
  - name: camera-A
    format: yolo
    images: ./repoA/images
    labels: ./repoB
    match: stem          # pair by filename; use `relpath` for parallel trees
    copy: ingest
vp sync --dry-run   # preview found / matched / unmatched / classes per source
vp sync             # ingest; idempotent, records per-asset provenance
vp sync --source camera-A   # sync just one source
vp sync --jobs 32   # concurrent transfers per source (remote defaults to 16+)

Sources can also live in object stores. Remote URIs go anywhere a local path would, and a target: lets copy mode land objects server-side in a content-addressed bucket without downloading them:

target: s3://my-bucket/datasets/factory-defects
sources:
  - name: camera-A
    format: yolo
    images: s3://my-bucket/raw/camera-a/images
    labels: s3://my-bucket/raw/camera-a/labels
    copy: copy

See Cloud Sync for credentials, copy modes, and streaming export.

Validate

Run validation:

vp validate

Strict mode treats missing annotations as errors:

vp validate --strict

Write a JSON validation report:

vp validate --report reports/validation.json

The current validator checks image readability, missing annotations, orphan labels, unknown classes, invalid boxes, boxes outside image bounds, duplicate exact assets, and split leakage.

Audit Label Health (vp audit)

vp validate catches labels that are invalid; vp audit finds labels that are valid but suspicious — the ones that usually turn out to be annotation mistakes:

vp audit

It reports duplicate boxes (the same object labeled twice), degenerate (tiny) boxes, boxes pinned to two or more image borders, whole-image boxes, extreme aspect-ratio outliers, rare classes, and dataset-level class imbalance.

Findings are advisory (exit code 0). Gate CI on them explicitly:

vp audit --fail-on-findings

Thresholds can be tuned per run (--min-box-px, --duplicate-iou, --max-aspect-ratio, --imbalance-ratio, --min-class-count) or persisted in visionpack.yaml:

validation:
  audit:
    min_box_px: 12
    duplicate_iou: 0.85

Like every pipeline command, vp audit --json prints a machine-readable envelope with per-code counts and the full findings list.

Show Statistics

Print a summary:

vp stats

Class distribution:

vp stats --by class

JSON output:

vp stats --json

Create Snapshots

Create a reproducible dataset snapshot:

vp snapshot create -m "initial import"

List snapshots:

vp snapshot list

Show one snapshot:

vp snapshot show v1

Snapshots store hashes for the manifest, assets, annotations, splits, and summary stats. They are written to .vp/snapshots/.

Lineage tags (which dataset trained this model?)

After a training run, stamp the snapshot it consumed:

vp snapshot tag v4 trained:run-812

Tags are free-form (the key:value convention is just a convention), show up in vp snapshot list/show, and can be removed with --remove. From the SDK, ds.snapshots_by_tag("trained:") lists every version any run trained on — so “which dataset produced this model?” is a lookup, not archaeology.

Diff Snapshots

Compare two snapshots:

vp diff v1 v2

JSON diff:

vp diff v1 v2 --json

The diff reports added and removed assets, added/removed/modified annotations, class changes, split changes, and before/after stats.

Distribution drift

--drift adds a class-distribution comparison: per-class object counts and distribution-share deltas (biggest movers first), plus KL and Jensen–Shannon divergence as single drift scores a CI job can threshold:

vp diff v1 v2 --drift
vp diff v1 v2 --drift --json   # adds a "drift" object to the diff payload

Because it derives from the stats frozen inside each snapshot, the drift between two versions is reproducible forever.

Export YOLO

Export the indexed dataset back to YOLO format:

vp export --format yolo --output exports/yolo-v1

The export writes:

exports/yolo-v1/
  images/
  labels/
  classes.txt
  data.yaml

Exports never duplicate bytes: local images are hardlinked from the content-addressed store, and cloud-backed images are written to a manifest.jsonl (image → object URI) for streaming instead of being downloaded. See Cloud Sync.

Segmentation projects export YOLO-seg polygon labels automatically (force either way with --seg / --no-seg), and --format masks writes semantic masks — 8-bit PNGs whose pixel value is the class index (0 = background, mapping recorded in classes.txt):

vp export --format yolo --output exports/seg-v1 --split      # YOLO-seg labels
vp export --format masks --output exports/masks-v1 --split   # class-index PNGs

Evaluate A Model (vp eval)

Score model predictions against the labels of a split’s set (the test set by default) — a locked split plus a snapshot makes the number a reproducible benchmark:

vp eval runs/predict/labels --format yolo         # Ultralytics save_txt output
vp eval predictions.json                          # vp-native or COCO JSON
vp eval predictions.json --split default --set val --json

Detection, segmentation, and keypoints report per-class AP@50, mAP@50, mAP@50-95, and precision/recall at a confidence threshold (--conf); classification reports accuracy, per-class precision/recall/F1, and a confusion matrix. Predictions reference images by asset id (what exports name files) or by original filename; unresolvable references are reported, never dropped silently.

Autolabel & The Annotation Queue

vp autolabel persists confident predictions as annotations. Model labels are recorded with source.type = "model", so they stay distinguishable from human labels. Only unlabeled assets are touched unless you pass --replace:

vp autolabel predictions.json --min-confidence 0.6

vp queue ranks images by how much a human label would help (active learning): unlabeled images first — ordered by model uncertainty when predictions are given — and with --include-labeled it audits existing labels for ground-truth / prediction disagreement (possible missing or stale labels):

vp queue --predictions predictions.json --limit 20
vp queue --predictions predictions.json --include-labeled --json

Together these close the model-in-the-loop cycle: export → train/predict → vp eval (measure) → vp autolabel (label the easy images) → vp queue (send the hard ones to humans) → re-train.

Pack Archive

Create a compressed archive package:

vp pack --profile archive

Or choose an output path:

vp pack --profile archive --output exports/archive/factory-defects.tar.zst

The archive includes:

  • visionpack.yaml
  • .vp/db/index.db
  • .vp/snapshots/*.json
  • content-addressed assets
  • pack.json with pack metadata and dataset stats

Python API

The supported programmatic surface is the SDK — the whole CLI workflow behind one Python class, with the same locking and result shapes (see the Python SDK page):

from visionpack.sdk import VisionPackClient

ds = VisionPackClient.open(".")
print(ds.name, len(ds))
ds.validate()
ds.export("./exports/yolo", format="yolo", split="default")

The lower-level visionpack.Dataset / Project handle stays available as ds.project for anything the facade doesn’t cover yet.

Current Limitations

  • segmentation metrics in vp eval use each polygon’s enclosing box (mask IoU is planned); YOLO-pose import/export and a dedicated keypoint importer are not implemented yet
  • vp annotate is scaffolded but not implemented yet
  • cloud sync (S3/GCS/Azure) is same-provider in v1 — cross-cloud transfer (S3↔GCS) and remote COCO/ImageFolder sync are planned; pack is local-only
  • the local index is SQLite (index.db); stats and exports stream records (flat RAM at scale), while validate, split, dedup, and the WebDataset pack still load the full set (next streaming pass — see ARCHITECTURE.md)

VisionPack is licensed under Apache-2.0. Built for the messy part of training computer-vision models.

This site uses Just the Docs, a documentation theme for Jekyll.