Patch-boundary artifacts
Independent block reconstruction leaves visible seams where neighboring volumetric patches meet along entire 2D faces.
Latent diffusion makes volumetric generation tractable, but its image autoencoder introduces a reconstruction bottleneck. We present VoxStruct3D, a voxel-space flow-matching framework that directly models full-resolution MRI volumes with a clean-data prediction objective.
Its Volumetric Voxel Generator (VVG) couples neighboring tokens through overlapping upsampling, refinement, and skip fusion. Structure First, Image Follows (SFIF) supplies an explicit anatomical prior using compact structure tokens, a structure-leading schedule, Patch-Aligned RoPE, and asymmetric attention.
On pathological and healthy T1-weighted brain MRI, VoxStruct3D delivers the strongest overall balance of feature-distribution alignment, sample diversity, and perceptual quality.
Direct generation preserves the original signal, but it exposes two problems that a plain patch-based model cannot solve on its own.
Independent block reconstruction leaves visible seams where neighboring volumetric patches meet along entire 2D faces.
Without an autoencoder’s compact abstraction, the model must recover global anatomy directly from a high-dimensional corrupted volume.
A shared model follows two synchronized clocks: the compact anatomy stream stays ahead while the image stream forms in the original voxel domain.
A DiT backbone handles global token interactions while an overlapping volumetric decoder lets neighboring tokens jointly reconstruct shared voxel regions.
Compact 3DINO-derived tokens act as an internal anatomical guide that develops ahead of, and only informs, the voxel stream.
Across pathological and healthy T1 brain MRI, VoxStruct3D achieves the best overall distribution alignment, diversity, and perceptual quality.
Additional T1 MRI volumes synthesized directly in voxel space. Each sample is shown from three orthogonal views.