Stanford University Multi-robot Systems Lab (MSL)
National University of Singapore Multi-Agent Robotic Motion Lab (MARMot)
Universidad de Zaragoza Perception-Oriented Control Team (POC), Universidad de Zaragoza

ScaffoldM3C A Multimodal Sequential Monte Carlo Framework for Generative Stable Construction Planning

Gadiel Sznaier Camps1Chengyang He2Guillaume Sartoretti2Eduardo Montijano3Mac Schwager1

1 Multi-robot Systems Lab (MSL), Stanford University

2 Multi-Agent Robotic Motion Lab (MARMot), National University of Singapore

3 Perception-Oriented Control Team (POC), Universidad de Zaragoza

225M
parameters, 4× smaller than BrickGPT
5–20×
faster inference
100%
collision-free build plans
2.3×
overall stability vs. BrickGPT-Scaffold

Overview

Given a starting build location, ScaffoldM3C auto-regressively proposes a set of next-block candidates, while a Sequential Monte Carlo search keeps several assembly sequences alive at once, populated by the candidate likelihoods.

Abstract

Autonomously constructing physically realizable 3D structures remains a significant challenge due to combinatorial action spaces, interchangeable components, equifinal assembly sequences, and strict stability requirements during construction. State-of-the-art methods fine-tune large language models for text-based generative construction. However, these approaches do not allow for multimodal (text, image, sketch) conditioning, overlook the practical role of scaffolding for stabilizing intermediate structures, and suffer from slow inference speeds.

We formulate construction as a probabilistic next-block generation task with multiple potential assembly actions and multiple potential task-conditioning modalities, and explicitly consider the utility of scaffolding by introducing an auxiliary scaffold block token. ScaffoldM3C (Scaffold Multimodal Monte Carlo) is a multimodal, lightweight, auto-regressive model for stable block-based construction that proposes a set of next-step candidate blocks. Leveraging these candidates, Sequential Monte Carlo maintains a population of possible assembly sequences, considering multiple, potentially different, assembly directions simultaneously.

We train on an extension of the StableText2Brick dataset with image conditioning prompts and scaffold-stabilized build sequences. ScaffoldM3C is 4× smaller than competing baselines, yielding a 5× to 20× inference speedup while achieving comparable construction quality and higher overall stability, and we demonstrate it through simulation and real-world robot assembly.

  • Multi-hypothesis assembly

    Candidate block proposals with placement likelihoods let SMC keep diverse, feasible build sequences alive instead of committing to one greedy rollout.

  • Scaffold tokens

    An auxiliary scaffold block (1×1 cross-section, adjustable height) resolves intermediate instability, such as overhangs, during generation.

  • Lightweight multimodal VLC

    A dual-stream Vision-Language-Construction transformer with interwoven cross-attention, trained from scratch on 553K text-, image- and multimodal-conditioned sequences.

Stability Criterion

A block is stable if it rests on the ground, or if the vertical projection p of its centroid c lies inside the 2D convex hull 𝓗 of the contact faces of the blocks directly beneath it. Scaffold blocks restore missing support so otherwise-unstable placements become feasible.

Stable Ground-supported A block resting on the ground is stable by direct ground contact.
Unstable Single support With one support, p falls outside 𝓗 and the block would tip over.
Stable Supports at both ends 𝓗 spans the gap between both contact faces, so p ∈ 𝓗.
Stable Scaffold-assisted An auxiliary scaffold block (gray) restores the missing support and enlarges 𝓗.
  • target block
  • support block
  • scaffold block
  • contact face
  • support region 𝓗
  • centroid c
  • ×projection p
  1. Grounded start. The first block b1 must contact the ground.
  2. No overlap. A newly placed block may not intersect existing ones.
  3. Stable on placement. Every block satisfies the criterion above the moment it is placed.

Combined four-panel clip: MP4 · GIF

Method

ScaffoldM3C maps a partial build sequence and a text and/or image prompt to a set of candidate next blocks (position, shape class, likelihood). Sequential Monte Carlo turns those candidates into full, feasible build plans.

Sequential Monte Carlo inference

  1. 1

    Predict

    Each of N particles queries the model for its own sequence-conditioned candidate set Ct,i.

  2. 2

    Measure & repair

    Collisions are repaired with a 2D minimum translation vector v; candidates are penalized by exp(−‖v‖).

  3. 3

    Reweight & resample

    Scores are normalized per particle, a block is sampled, and surviving particles are importance-resampled.

  4. 4

    Select

    Sequences ending in EOS are harvested; the longest, highest-weight completed plan is returned.

Architecture. Block positions are lifted with Fourier features and concatenated with class embeddings; prompts are encoded by a 4-bit quantized Gemma 3. Interwoven causal and cross-attention layers feed a windowed attention layer and three cascaded heads that predict candidate positions, classes and likelihoods.

A multimodal, stability-aware dataset

We build on StableText2Brick. Each original sequence is replayed block by block; whenever a placement violates our stability criterion, the minimum number of scaffold blocks needed to support it is inserted. Structure geometry and target-block order are preserved, so every intermediate assembly is stable.

Each structure then receives three kinds of visual condition:

  • Synthetic: the build is simulated and rendered from a sampled viewpoint.
  • Realistic: the object in a natural scene, from a Qwen text-to-image model.
  • Abstract: a simplified artistic interpretation from the same model.

Every sequence expands into 13 condition–sequence pairs (5 text-only, 3 image-only, 5 multimodal), for 553K conditioned assembly sequences in total.

Results

Evaluated without rollbacks on 479 text and 288 image test prompts against BrickGPT, a 1B-parameter LLaMA-3.2 model fine-tuned for brick assembly, and a scaffold-aware variant we fine-tuned (BrickGPT-Scaffold).

26.7% fully stable & feasible structures (text prompts, SMC) vs. 11.3% for BrickGPT-Scaffold
100% collision-free rate, up from 90.2% for BrickGPT-Scaffold
1.8s mean SMC inference time vs. 11.5–64.5 s for BrickGPT variants
0.84 DINOv2 similarity to ground truth, the highest among all methods
Average results without rollbacks. ↑ higher is better, ↓ lower is better, ≈ closer to ground truth is better.
Method soverall ↑ efeasible ↑ erate ↑ sfinal ↑ agt ↑ blocks ≈ scaffolds ≈ time (s) ↓
Text prompts
Ground truth100.00%100.00%100.00%1.001.00111.028.0—
BrickGPT w/o rej.2.73%2.31%28.09%0.760.82106.20.011.46
BrickGPT w/ rej.3.13%3.13%100.00%0.760.82110.90.013.86
BrickGPT-Scaffold11.48%11.27%90.19%0.920.83195.142.464.51
Ours (Top-K)20.04%20.04%100.00%0.820.8189.621.30.58
Ours (SMC)26.72%26.72%100.00%0.840.84126.929.21.79
Image prompts (BrickGPT variants cannot take image input)
Ground truth100.00%100.00%100.00%1.001.00110.928.7—
Ours (Top-K)18.40%18.40%100.00%0.810.8295.023.50.62
Ours (SMC)20.49%20.49%100.00%0.830.84128.230.22.13

BrickGPT-Scaffold's higher per-block sfinal comes from over-generating inherently stable scaffold tokens (about 1.45× more than ours); it rarely yields a fully stable structure.

Qualitative results on text and image prompts from the test set. SMC produces structures visually comparable to BrickGPT with fewer artifacts than Top-K, and distributes scaffolds where they are needed.

Robot Hardware Experiments

Generated plans are built autonomously by a 6-DOF robotic arm from 3D-printed, LEGO-style blocks. Each scaffold token is physically instantiated as a stack of singleton scaffold blocks, which hold up overhangs such as the prow of the boat, the bumpers of the car and the shelves of the bookshelf.

Real world · 32× speed · wrist + side camera
Simulated plan

Prompt Image only: a photo of a car (left)

Car image prompt and four views of the predicted structure
Finished builds and the nine block types used. Yellow scaffold blocks support overhangs that would otherwise be unstable.

Ablations & Analysis

SMC lets users trade inference time for structural quality through the particle population size and the candidate and class temperatures.

Population size. Larger populations yield more complete, coherent structures, with diminishing returns beyond 160 sequences.
SMC population size (24 test samples)
Nsoverall ↑sfinal ↑time (s) ↓
120.83%0.800.42
420.83%0.810.56
1020.83%0.850.89
2025.00%0.881.60
4020.83%0.882.64
8033.33%0.914.93
16029.17%0.9311.09
32033.33%0.9122.89

Highlighted rows: the best quality/time trade-off (N = 20–40).

Candidate vs. class temperature. Stability is highly sensitive to the class temperature; raising the candidate temperature broadens exploration and generally improves stability.
SMC belief state after 12 steps generating a table (4 particles). Each particle holds its own candidate set, so alternative placements are explored simultaneously.
Failure modes: block misalignment, early termination, prompt confusion, and missing blocks or scaffolding, most often downstream of an initial misalignment.

Resources

Paper arXiv Preprint Full paper with appendix arXiv:2610.00487 →
Code GitHub Repository Open-source model, dataset tools and SMC inference Coming soon

BibTeX

Coming soon

The citation entry will be posted here once the arXiv version is available.