Editable graphic design
PlayGround mascot holding a paintbrush, from rectangle (1).pdf

PlayGround

Progressive Layout Generationwith Render-Grounded Planning

Pipeline Overview

TL;DR

We propose PlayGround, a progressive layout generation framework that adaptively adjusts the design workflow based on the current rendered canvas. Under our six-criterion VLM evaluation, PlayGround performs favorably against existing methods on both Crello and LICA, including zero-shot evaluation on LICA.

Abstract

Generating editable graphic designs requires arranging visual and textual elements into a coherent composition while preserving their content and editability. Although recent MLLM-based approaches employ iterative generation and visual refinement to reflect the real-world design workflows, their high-level plans are predetermined prior to execution, limiting their adaptability as the design evolves. To address this limitation, we propose PlayGround, a progressive layout generation framework that dynamically adapts to the current rendered canvas. PlayGround establishes a progressive design loop comprising a Progressive Planner, a Layout Generator, and a Visual Verifier, guided by a Structure-Aware Retriever. Given an initial set of unordered elements, the Structure-Aware Retriever initially infers a coarse layout structure to retrieve structurally similar reference designs. Subsequently, at each turn, the Progressive Planner examines the current rendered canvas alongside the remaining elements, forms the next functional group, and plans its composition. The Layout Generator then predicts editable layout attributes, while the Visual Verifier evaluates rendered candidates and provides feedback when refinement is necessary. To train the Progressive Planner with step-wise supervision, we construct a two-stage data generation pipeline that extracts functional groups and turn-level planning trajectories from completed designs. Experiments on Crello and LICA demonstrate the effectiveness of PlayGround over existing methods. Our code and weights will be publicly released.

Data Construction

From completed designs to functional groups and group-wise turn-level planning supervision. The two-stage pipeline produces group regions, cumulative canvas states, and group-wise turn-level plans, which are used to train the Structure-Aware Retriever, Progressive Planner, and Layout Generator.

Two-stage data construction pipeline
Two-stage data construction pipeline
Stage 1 constructs functional groups and their execution order. Stage 2 derives a turn-level plan for each group from the before-and-after rendered states.
STAGE 1

Functional grouping & execution order

An MLLM annotator groups elements that should be planned and positioned together, assigns functional roles, and determines a dependency-aware execution order. Group regions and cumulative canvas states are then derived.

STAGE 2

Turn-level planning annotation

With grouping and execution order fixed, before-and-after renders provide visual evidence for each group’s placement. The annotator produces grouping, execution, and layout rationales together with a qualitative layout instruction.

PlayGround

Overall pipeline
Complete PlayGround pipeline with retrieval, progressive planning, layout generation, and visual verification
Select a module to follow the pipeline
The accepted canvas and remaining elements define the input state for the following turn.
01

Structure-Aware Retriever

Global composition guidance · Once before generation

The Retriever infers a coarse layout structure from the input elements as functional group role–region pairs. It retrieves five references using role-aware spatial similarity, canvas aspect ratio, and text–image content counts. These references remain fixed throughout generation.

Query and retrieved references
Query and retrieved references
The retrieved designs exhibit similar coarse layout structures and canvas geometry.
02

Progressive Planner

Dynamic grouping · Render-grounded composition

The Planner observes the current canvas and remaining elements, forms a functional group, and plans its composition. In the example below, it uses the available upper-left space for text, while the Static Planner, which must produce the entire plan at once and therefore imagine the intermediate canvases, targets a central placard already occupied by photographs.

Progressive versus static planning
Progressive versus static planning
The same current group, different planning context. Our Planner adapts to the rendered canvas rather than relying on a fixed initial plan.
03

Layout Generator

From a qualitative plan to editable attributes

The Layout Generator predicts element-level attributes for the current group while keeping previously placed elements unchanged. It generates multiple candidate layouts because a group-level plan can admit different rendered outcomes.

Bounding boxesGroup-local stackingText attributes
Five actual candidate renders cropped from the supplied pipeline figure

Five candidate renders from the pipeline overview.

04

Visual Verifier

Inspect rendered candidates · Accept or retry

The Verifier selects a rendered candidate and either accepts it or returns a revised layout instruction for regeneration. Up to three verification rounds are allowed; if none is accepted, it makes a final selection among the per-round finalists.

Accept
Update the canvas and remaining elements.
Retry with feedback
Revise the instruction. Regenerate the same group.

Results

Quantitative comparison

Table 1. Six VLM-based criteria (mean ± sample standard deviation) and two rule-based metrics. VLM scores range from 1 to 5; higher is better. T.Occ. (text occlusion) is the fraction of text glyph pixels hidden by elements stacked above them; Rea. (text unreadability) measures how visually busy the background beneath the text is. Lower is better for both.

Crello quantitative results from Table 1 of the supplied manuscript.
MethodVLM-based metrics ↑Rule-based metrics ↓
(i)
Composition
(ii)
Typography
(iii)
Elements
(iv)
Color
(v)
Style
(vi)
Intention
T.Occ. Rea.
Bold and underline mark the best and second-best means among the evaluated methods, excluding GT. VFLM places text on a fixed ground-truth background. Text-free sources are excluded from the typography aggregate and from all VFLM aggregates.
(i) Layout composition rationality(ii) Typography legibility and hierarchy(iii) Visible element presentation(iv) Color harmony and contrast(v) Visual style coherence(vi) Design intention alignment

Qualitative comparisons

Qualitative comparison · Crello
Qualitative comparison · Crello
From the same supplied elements. We compare rendered outputs with the corresponding ground-truth designs. Columns follow the source figure: GT, FlexDM, LaDeCo, PosterCopilot, VFLM, and PlayGround.

Extension to
Intention-to-Design

Users may have a design intention without the visual and textual elements needed to realize it. We extend PlayGround to this setting by prepending an element-generation stage.

GPT-5.6 Luna derives the textual elements and visual-element descriptions. Qwen-Image-2.1 generates the corresponding visual elements, which are then composed by PlayGround without modifying the core pipeline.

These examples on Crello demonstrate the modular extensibility of PlayGround as a downstream composition module, from element-to-design to intention-to-design generation.

Design intention
Element generation
PlayGround
Editable design
Intention-to-design · Qualitative results
Intention-to-design · Qualitative results
Qualitative examples of PlayGround applied to elements generated from a design intention.

Citation

BIBTEX · DRAFT
% Add final publication metadata before publishing.
@misc{playground,
  title = {{PlayGround}: Progressive Layout Generation
           with Render-Grounded Planning},
  author = {Lee, Jaeho and Heo, Junhwan and Park, Junghyun and Moon, Wonjun
            and Yang, Eunju and Jang, Seungho and Kim, Seongchan and Kim, Seungryong},
  note = {Manuscript}
}

Author names are provided above. Final publication metadata will be added when available; this citation remains a draft.