Arcee AI: Spotlight

Arcee AIID: arcee-ai/spotlight

Spotlight is a 7‑billion‑parameter vision‑language model derived from Qwen 2.5‑VL and fine‑tuned by Arcee AI for tight image‑text grounding tasks. It offers a 32 k‑token context window, enabling rich multimodal conversations that combine lengthy documents with one or more images. Training emphasized fast inference on consumer GPUs while retaining strong captioning, visual‐question‑answering, and diagram‑analysis accuracy. As a result, Spotlight slots neatly into agent workflows where screenshots, charts or UI mock‑ups need to be interpreted on the fly. Early benchmarks show it matching or out‑scoring larger VLMs such as LLaVA‑1.6 13 B on popular VQA and POPE alignment tests.

Pricing per 1M Tokens

Input (Prompt)$0.18
Output (Completion)$0.18
Cache ReadFree
Cache WriteFree
ImageN/A

Specifications

Context Length131K
Max Output Tokens66K
Input ModalitiesImage + Text
Output ModalitiesText
TokenizerOther
Instruct TypeN/A
Top Provider Context131K
Top Provider Max Output66K
ModeratedNo

More from Arcee AI

Last updated: March 23, 2026

First tracked: March 23, 2026