ScaffoldM3C A Multimodal Sequential Monte Carlo Framework for Generative Stable Construction Planning
- 225M
- parameters, 4× smaller than BrickGPT
- 5–20×
- faster inference
- 100%
- collision-free build plans
- 2.3×
- overall stability vs. BrickGPT-Scaffold
Overview
Abstract
Autonomously constructing physically realizable 3D structures remains a significant challenge due to combinatorial action spaces, interchangeable components, equifinal assembly sequences, and strict stability requirements during construction. State-of-the-art methods fine-tune large language models for text-based generative construction. However, these approaches do not allow for multimodal (text, image, sketch) conditioning, overlook the practical role of scaffolding for stabilizing intermediate structures, and suffer from slow inference speeds.
We formulate construction as a probabilistic next-block generation task with multiple potential assembly actions and multiple potential task-conditioning modalities, and explicitly consider the utility of scaffolding by introducing an auxiliary scaffold block token. ScaffoldM3C (Scaffold Multimodal Monte Carlo) is a multimodal, lightweight, auto-regressive model for stable block-based construction that proposes a set of next-step candidate blocks. Leveraging these candidates, Sequential Monte Carlo maintains a population of possible assembly sequences, considering multiple, potentially different, assembly directions simultaneously.
We train on an extension of the StableText2Brick dataset with image conditioning prompts and scaffold-stabilized build sequences. ScaffoldM3C is 4× smaller than competing baselines, yielding a 5× to 20× inference speedup while achieving comparable construction quality and higher overall stability, and we demonstrate it through simulation and real-world robot assembly.
-
Multi-hypothesis assembly
Candidate block proposals with placement likelihoods let SMC keep diverse, feasible build sequences alive instead of committing to one greedy rollout.
-
Scaffold tokens
An auxiliary scaffold block (1×1 cross-section, adjustable height) resolves intermediate instability, such as overhangs, during generation.
-
Lightweight multimodal VLC
A dual-stream Vision-Language-Construction transformer with interwoven cross-attention, trained from scratch on 553K text-, image- and multimodal-conditioned sequences.
Stability Criterion
A block is stable if it rests on the ground, or if the vertical projection p of its centroid c lies inside the 2D convex hull 𝓗 of the contact faces of the blocks directly beneath it. Scaffold blocks restore missing support so otherwise-unstable placements become feasible.
- target block
- support block
- scaffold block
- contact face
- support region 𝓗
- centroid c
- ×projection p
- Grounded start. The first block b1 must contact the ground.
- No overlap. A newly placed block may not intersect existing ones.
- Stable on placement. Every block satisfies the criterion above the moment it is placed.
Method
ScaffoldM3C maps a partial build sequence and a text and/or image prompt to a set of candidate next blocks (position, shape class, likelihood). Sequential Monte Carlo turns those candidates into full, feasible build plans.
Sequential Monte Carlo inference
-
1
Predict
Each of N particles queries the model for its own sequence-conditioned candidate set Ct,i.
-
2
Measure & repair
Collisions are repaired with a 2D minimum translation vector v; candidates are penalized by exp(−‖v‖).
-
3
Reweight & resample
Scores are normalized per particle, a block is sampled, and surviving particles are importance-resampled.
-
4
Select
Sequences ending in EOS are harvested; the longest, highest-weight completed plan is returned.
A multimodal, stability-aware dataset
We build on StableText2Brick. Each original sequence is replayed block by block; whenever a placement violates our stability criterion, the minimum number of scaffold blocks needed to support it is inserted. Structure geometry and target-block order are preserved, so every intermediate assembly is stable.
Each structure then receives three kinds of visual condition:
- Synthetic: the build is simulated and rendered from a sampled viewpoint.
- Realistic: the object in a natural scene, from a Qwen text-to-image model.
- Abstract: a simplified artistic interpretation from the same model.
Every sequence expands into 13 condition–sequence pairs (5 text-only, 3 image-only, 5 multimodal), for 553K conditioned assembly sequences in total.
Results
Evaluated without rollbacks on 479 text and 288 image test prompts against BrickGPT, a 1B-parameter LLaMA-3.2 model fine-tuned for brick assembly, and a scaffold-aware variant we fine-tuned (BrickGPT-Scaffold).
| Method | soverall ↑ | efeasible ↑ | erate ↑ | sfinal ↑ | agt ↑ | blocks ≈ | scaffolds ≈ | time (s) ↓ |
|---|---|---|---|---|---|---|---|---|
| Text prompts | ||||||||
| Ground truth | 100.00% | 100.00% | 100.00% | 1.00 | 1.00 | 111.0 | 28.0 | — |
| BrickGPT w/o rej. | 2.73% | 2.31% | 28.09% | 0.76 | 0.82 | 106.2 | 0.0 | 11.46 |
| BrickGPT w/ rej. | 3.13% | 3.13% | 100.00% | 0.76 | 0.82 | 110.9 | 0.0 | 13.86 |
| BrickGPT-Scaffold | 11.48% | 11.27% | 90.19% | 0.92 | 0.83 | 195.1 | 42.4 | 64.51 |
| Ours (Top-K) | 20.04% | 20.04% | 100.00% | 0.82 | 0.81 | 89.6 | 21.3 | 0.58 |
| Ours (SMC) | 26.72% | 26.72% | 100.00% | 0.84 | 0.84 | 126.9 | 29.2 | 1.79 |
| Image prompts (BrickGPT variants cannot take image input) | ||||||||
| Ground truth | 100.00% | 100.00% | 100.00% | 1.00 | 1.00 | 110.9 | 28.7 | — |
| Ours (Top-K) | 18.40% | 18.40% | 100.00% | 0.81 | 0.82 | 95.0 | 23.5 | 0.62 |
| Ours (SMC) | 20.49% | 20.49% | 100.00% | 0.83 | 0.84 | 128.2 | 30.2 | 2.13 |
BrickGPT-Scaffold's higher per-block sfinal comes from over-generating inherently stable scaffold tokens (about 1.45× more than ours); it rarely yields a fully stable structure.
Robot Hardware Experiments
Generated plans are built autonomously by a 6-DOF robotic arm from 3D-printed, LEGO-style blocks. Each scaffold token is physically instantiated as a stack of singleton scaffold blocks, which hold up overhangs such as the prow of the boat, the bumpers of the car and the shelves of the bookshelf.
Ablations & Analysis
SMC lets users trade inference time for structural quality through the particle population size and the candidate and class temperatures.
| N | soverall ↑ | sfinal ↑ | time (s) ↓ |
|---|---|---|---|
| 1 | 20.83% | 0.80 | 0.42 |
| 4 | 20.83% | 0.81 | 0.56 |
| 10 | 20.83% | 0.85 | 0.89 |
| 20 | 25.00% | 0.88 | 1.60 |
| 40 | 20.83% | 0.88 | 2.64 |
| 80 | 33.33% | 0.91 | 4.93 |
| 160 | 29.17% | 0.93 | 11.09 |
| 320 | 33.33% | 0.91 | 22.89 |
Highlighted rows: the best quality/time trade-off (N = 20–40).
Resources
BibTeX
Coming soonThe citation entry will be posted here once the arXiv version is available.
Gadiel Sznaier Camps1
Chengyang He2
Guillaume Sartoretti2
Eduardo Montijano3
Mac Schwager1