Presentation
49. SAWNA: Space-Aware Text to Image Generation
DescriptionSAWNA tackles layout-sensitive text-to-image generation by treating user-specified empty regions as first-class constraints. Bounding-box masks are blurred and injected as mean-shifted, inert noise into the frozen Stable Diffusion latent, suppressing synthesis inside reserved areas while preserving diversity and quality elsewhere. This simple training-free modification supports workflows that require precise layout fidelity, including advertising (e.g., space for logos or headlines), UI design (e.g., button placement), and animation pre-production (e.g., speech bubbles, subtitles, or motion overlays).
Experiments show that SAWNA outperforms layout-aware baselines like GLIGEN and in-painting pipelines, both of which struggle to maintain truly empty regions without introducing artifacts or incoherence. In contrast, SAWNA yields clean, editable space while producing semantically rich images across the remaining canvas.
This makes it especially suitable for design-critical applications where reserved regions are integral to downstream compositing or storytelling.
Experiments show that SAWNA outperforms layout-aware baselines like GLIGEN and in-painting pipelines, both of which struggle to maintain truly empty regions without introducing artifacts or incoherence. In contrast, SAWNA yields clean, editable space while producing semantically rich images across the remaining canvas.
This makes it especially suitable for design-critical applications where reserved regions are integral to downstream compositing or storytelling.

Event Type
Poster
TimeThursday, 14 August 20259:00am - 5:30pm PDT
LocationWest Building, Level 2, Outside Room 219
Session TimeSunday, 10 August 20259:00am - 5:30pm PDTMonday, 11 August 20259:00am - 5:30pm PDTTuesday, 12 August 20259:00am - 5:30pm PDTWednesday, 13 August 20259:00am - 5:30pm PDTThursday, 14 August 20259:00am - 5:30pm PDT
LocationWest Building, Level 2, Outside Room 219
Similar Presentations
