Vision catalogue · Verified local path

Florence-2 Base local guide

Compact general vision model for captioning, OCR, detection, segmentation and grounding.

Choose an app

Start with Recommended. No terminal commands are shown.

Compare vision models

What it does

Image CaptioningOcrObject DetectionSegmentation

The compact base checkpoint is suitable for modest local hardware through Transformers or optimized runtimes.

Hardware figures are practical entry floors, not performance guarantees. Resolution, duration, precision, offloading and runtime versions can materially change memory use.

Strengths

  • Many vision tasks
  • Compact checkpoint
  • MIT license

Limits to know

  • Task prompts require exact formatting