EMNLP 2026

Controllable Image Captioning with Prompt-Conditioned Scene Rewards

1Graduate School of Artificial Intelligence, POSTECH
2Department of Computer Science and Engineering, POSTECH
*Equal contribution

Code will be added when available.

Comparison showing FoCUS emphasizing surfboard attributes while zero-shot captioning gives a generic scene description
Zero-shot generation defaults to a generic scene description, whereas FoCUS follows the prompt and emphasizes object attributes.

Abstract

Large Vision-Language Models produce fluent image descriptions but offer limited semantic control: users cannot reliably specify whether captions should emphasize attributes, relations, or particular image regions. We present Fine-grained Captioning Control Using Scene-rewards (FoCUS), a controllable image captioning method that lets users steer captions toward specific semantic emphases through natural-language control prompts. The core idea is a prompt-conditioned control objective based on scene-graph-aligned component scores. Generated captions are parsed and aligned to scene-graph components—objects, attributes, and relations—and these components are differentially weighted, including negative weights, according to the requested emphasis. We optimize this objective with GRPO and further improve its reliability through a stricter object validity threshold and reasoning-based verification for attribute and relation scoring. To evaluate controllability, we introduce Semantic Control and Precision Evaluation (SCoPE), a benchmark with contrastive Include/Avoid constraints for measuring both target content coverage and out-of-scope suppression. Experiments on two VLM backbones show that FoCUS consistently improves controllability and fine-grained caption quality without degrading general caption performance.

Motivation

Modern vision-language models generate fluent and detailed captions, but they usually produce a single “best overall” description. For the same image, users may instead want a caption centered on visual attributes, spatial relations, foreground objects, or background context.

Limited semantic control

Current captioners provide few reliable ways to decide what information a caption should emphasize.

Prompting is inconsistent

Natural-language prompts often drift toward generic, high-probability descriptions and include details outside the requested focus.

Structured inputs are less natural

Length tokens, regions, boxes, and formal scene graphs can provide control, but are less convenient than language-based interaction.

Method

FoCUS decomposes each generated caption into scene-graph components and aggregates object, attribute, and relation scores using prompt-specific signed weights. Positive weights reward requested content, while negative weights suppress off-scope content. The resulting objective is optimized with a two-stage SFT + GRPO pipeline.

FoCUS pipeline with an image and control prompt, scene-graph matching, component scoring, and prompt-conditioned reward aggregation
Overview of the FoCUS pipeline. The generated caption is aligned with scene-graph annotations and scored at the object, attribute, and relation levels.

Scene-aligned scoring

Strict object matching and reasoning-based attribute/relation verification make the training reward more reliable.

Prompt-conditioned rewards

Five prompt categories—General, Attribute, Relation, Foreground, and Background—receive distinct signed reward weights.

GRPO optimization

Multiple candidate captions provide group-relative advantages for stable, critic-free policy optimization.

SCoPE Benchmark

Semantic Control and Precision Evaluation (SCoPE) evaluates controllable captioning through category-specific Include and Avoid lists. It measures whether a caption covers the requested facts, suppresses content outside the requested focus, and remains factually consistent.

SCoPE construction and evaluation pipeline using category-specific atomic fact Include and Avoid lists
SCoPE evaluates target coverage, off-scope suppression, and factual consistency with contrastive atomic-fact lists.
189Curated images
4Control categories
CoverageRequested facts
AdherenceOff-scope suppression
FaithfulnessFactual consistency

Results

Across both VLM backbones, SFT + FoCUS achieves the best overall SCoPE score and improves fine-grained factual alignment under CompreCap.

Table 1 comparing FoCUS with zero-shot, SFT, CLIP, and CompreCap baselines on SCoPE and CompreCap metrics
Table 1. Performance comparison on SCoPE and CompreCap across two VLM backbones.

Qualitative Examples

Original excerpts from Figure 9. The same image is captioned under four different control prompts.

Red text follows the highlights in Figure 9.

A woman in a black dress looking at a smartphone in a warmly lit indoor setting
Example image used in Figure 9.

Attribute-focused

Zero-shot

The image shows a person standing indoors, likely in a dimly lit room. They are wearing a black dress with thin straps and a headband. The person has large hoop earrings and is holding a smartphone, seemingly focused on it. ... They are also wearing a necklace and a sparkly, possibly sequined, jacket draped over their shoulders. ... The person's expression is engaged, possibly looking at the phone screen.

FoCUS

A woman with dark brown hair pulled back and light skin shows small freckles on her face and wears a black sleeveless dress with thin straps and a smooth, matte fabric; she has large black circular hoop earrings and a thin gold-colored chain necklace. ... A glowing neon ring with a blue-purple light is visible. ... The wall is painted pale pink with a red circular printed motif.

Relation-focused

Zero-shot

The image shows a person standing indoors, likely in a dimly lit room. The individual is focused on a smartphone they are holding in their hands. They are wearing a black dress with thin straps and a headband. Around their neck, they have a necklace and large hoop earrings. A glowing neon ring is visible around their waist, adding a colorful accent to their outfit. ... In the background, there is a wall decorated with string lights, creating a warm ambiance. ... The room has a cozy, festive feel, possibly decorated for an event or celebration.

FoCUS

A woman stands in the foreground with a vertical support pole behind her head, a string of lights to her left and a framed picture on the wall behind and slightly to her left; a television sits behind and to her right, while a jacket draped over her shoulder hangs beside her right side and a glowing neon band encircles her waist between her torso and hands, with her hands positioned on a small rectangular device held in front of her face.

Foreground-focused

Zero-shot

The image shows a person standing indoors, likely in a dimly lit room. The individual is focused on a smartphone they are holding in their hands. ... In the background, there is a wall decorated with string lights, creating a warm ambiance. A framed picture hangs on the wall, and a television screen is visible, displaying some content. ... possibly decorated for an event or celebration.

FoCUS

A woman in a black sleeveless dress and large black hoop earrings stands holding a smartphone, her face illuminated by the phone's screen while a glowing purple ring accentuates her waist.

Background-focused

Zero-shot

The image shows a person standing indoors, likely in a dimly lit room. The individual is wearing a black dress and has large hoop earrings. They are focused on a smartphone, holding it with both hands. The person has a headband and is accessorized with a necklace and a bracelet. In the background, there are string lights creating a warm ambiance. ...

FoCUS

An indoor residential space with a painted wall featuring a circular red emblem and a framed picture, a blue-painted vertical support post, ... and a glimpse of a countertop and countertop lights, all illuminated by warm artificial lighting that casts soft shadows across the room.

BibTeX

@misc{hyun2026controllable,
  title={Controllable Image Captioning with Prompt-Conditioned Scene Rewards},
  author={Jongyeop Hyun and Taeyoung Kim and Hyounghun Kim},
  year={2026},
  eprint={2609.00709},
  archivePrefix={arXiv},
  primaryClass={cs.CV},
  url={https://arxiv.org/abs/2609.00709},
}