arXiv is now an independent nonprofit! Learn more
License: CC BY 4.0
arXiv:2609.30234v1 [cs.CV] 24 Sep 2026

OmniFabric: Coherent UV Space Texture Synthesis for 3D Garment Reconstruction

Conference: SIGGRAPH Asia 2026 Conference Papers; December 01–04, 2026; Kuala Lumpur, MalaysiaSIGGRAPH Asia 2026 Conference Papers (SA Conference Papers ’26), December 01–04, 2026, Kuala Lumpur, MalaysiaDOI: 10.1145/3829340.3842275ISBN: 979-8-4007-2842-6/2026/121837CCS: Computing methodologies Appearance and texture representations
Ding-Jiun Huang Affiliation: Carnegie Mellon University, United States of America email: djhuang322@gmail.com , Yuanhao Wang Affiliation: University of Washington, United States of America email: yuanhao4@cs.washington.edu , Cheng Zhang Affiliation: Texas A&M University, United States of America email: chzhang@tamu.edu , Hugo Bertiche Affiliation: Google, United States of America email: hbertiche@google.com , Alexandru-Eugen Ichim Affiliation: Google, Switzerland email: alexichim@google.com , Thabo Beeler Affiliation: Google, Switzerland email: tbeeler@google.com and Fernando De la Torre Affiliation: Carnegie Mellon University, United States of America email: ftorre@andrew.cmu.edu
© cc

Four example pairs showing input images and simulation results.

Figure 1. Image-based 3D garment texture synthesis. Given a single in-the-wild clothing image, OmniFabric synthesizes high-quality, coherent fabric textures directly on garment sewing patterns. OmniFabric supports a wide range of textures and patterns, preserving fine details while maintaining global alignment with the input image. The resulting textured sewing patterns can be seamlessly simulated into 3D garments. Four example pairs showing input images and simulation results.
Abstract.

Automated generation of production-ready 3D garment assets from a single image is a central challenge in digital content creation. While recent generative models have significantly advanced 3D geometry reconstruction, synthesizing high-quality textures remains a bottleneck. Existing methods often bake environmental illumination and shadows directly into the texture map, or they fail to maintain global structural coherence, making the resulting assets unusable for physical simulation and relighting. In this work, we introduce OmniFabric, a novel approach that synthesizes globally coherent texture maps directly within the 2D sewing pattern space. Given a single reference image, our pipeline utilizes an estimated 3D mesh and generative priors of powerful Vision-Language Models (VLM) to establish a complete but coarse texture initialization across the unwrapped sewing patterns. We then leverage a specialized diffusion transformer, trained via an automated synthetic data engine and conditioned on 3D positional features, to refine this initialization directly in the canonical UV domain. This effectively removes distortion and baked-in artifacts to extract a clean and normalized texture map that preserves the original garment design. Extensive experiments demonstrate that OmniFabric significantly outperforms state-of-the-art baselines, yielding photorealistic 3D garments with high-quality textures. Project page: https://humansensinglab.github.io/OmniFabric

Keywords: 
Texture Synthesis, 3D Garment Reconstruction, Garment Sewing Patterns
††cc-license: by

1. Introduction

Creating high-fidelity, simulation-ready 3D garments is essential for modern digital production, with applications spanning gaming (Habermann et al., 2021), e-commerce (Hwangbo et al., 2020; Kim and LaBat, 2013), and cinematic visual effects (Hughes et al., 2007). While traditional asset creation relies on labor-intensive manual modeling, there is growing interest in automating this process to synthesize assets directly from images. Despite recent strides in reconstructing accurate 3D garment geometry (Bian et al., 2025; Nakayama et al., 2025; Sarafianos et al., 2025; He et al., 2024; Li et al., 2025b; Rong et al., 2025; Wang et al., 2025), generating production-ready textures from a single observation remains an open problem.

Garment texture synthesis is uniquely difficult due to the complex dynamics of fabric, which frequently produce deep folds, self-occlusions, and extreme pose variations (see Figure 1). To be truly usable in downstream applications, a texture must not only look realistic in a static view but also be free from baked-in lighting and geometric distortion, capturing the garment’s intrinsic albedo. However, current state-of-the-art 3D texturing frameworks (Richardson et al., 2023; Chen et al., 2023a; Zeng et al., 2024; Zhao et al., 2025; Bensadoun et al., 2024) typically rely on multi-view diffusion priors to iteratively “paint” assets via direct projection onto draped 3D geometry. Consequently, unwrapping these folded surfaces inevitably causes large geometric distortion and occlusion. Furthermore, these models fail to disentangle intrinsic material properties from environmental illumination, baking transient lighting artifacts, such as shadows and geometry-induced wrinkles, directly into the texture map.

To circumvent these issues, recent work such as FabricDiffusion (Zhang et al., 2024) has explored texture synthesis directly within the regularized, flat 2D sewing pattern domain, successfully yielding simulation-ready texture maps. However, FabricDiffusion focuses exclusively on local texture synthesis by extracting tileable material patches from the source image, ignoring the global texture layout. Consequently, it struggles at reconstructing asymmetric designs, large-scale prints, or irregular spatial patterns. We argue that synthesizing a globally coherent texture map that faithfully reflects the visual identity of the input image requires leveraging the complete context of the reference observation, rather than relying on isolated patches. This presents a fundamental challenge:

How do we establish pixel-level correspondences between
a reference image and 2D sewing panels while transferring textures
that are free from baked-in artifacts and geometric distortions?

In this paper, we propose OmniFabric to bridge this gap. We first employ a powerful image generation model (Google, 2025) to repose the input clothing into a canonical A-pose, aligning its silhouette with the frontal render of the garment mesh simulated from the predicted sewing patterns. We further harness the temporal and spatial consistency of large-scale video generation models (Google DeepMind, 2025) to hallucinate coherent, multi-view observations of this A-pose garment. Projecting these globally consistent multi-view images onto the unwrapped sewing patterns establishes a complete, albeit coarse, texture initialization that preserves the garment’s global design. Because this initial projection inevitably bakes visual artifacts such as geometric distortions and physics-induced wrinkles into the texture map, we introduce a subsequent refinement stage operating entirely within the sewing pattern space. We design an automatic data engine that curates an extensive and highly diverse synthetic dataset of textured sewing patterns, paired with simulated texture initializations containing baked-in artifacts and distortions. With the curated dataset, we train a specialized Diffusion Transformer (DiT) that acts as a texture normalizer in the 2D sewing pattern space. This model rectifies distortions, and strips away transient illumination to produce a clean and normalized texture map. Finally, we conduct comprehensive experiments to demonstrate the superiority of OmniFabric over existing 3D texturing baselines. In summary, the key contributions of our work are:

  • •

    OmniFabric, a novel framework for synthesizing high-fidelity, globally coherent textures from a single in-the-wild image for 3D garment reconstruction.

  • •

    An automated data engine for creating realistic, diverse textured sewing patterns at scale, providing training data for learning complex garment textures.

  • •

    A strategy of using a canonical 3D mesh as spatial anchor to map image pixels to sewing pattern UV space, enabling globally aligned texture synthesis.

  • •

    A coarse-to-fine texture generation pipeline that initializes globally aligned textures, and further rectifies distortion and baked-in artifacts with a DiT-based UV space normalization network.

Remark.

Generating 3D garment textures is not a new topic. While many prior works mainly focus on directly texturing 3D garment meshes, they often do not scale well to large collections of in-the-wild images and remain weakly connected to real-world garment production pipelines. We argue that fabric texture synthesis should be grounded in the UV space of sewing patterns. This enables the model to maintain strong global coherence while remaining closely aligned with how garments are constructed in practice.

2. Related Work

2.1. 3D Garment Reconstruction

Early approaches for garment reconstruction primarily relied on implicit functions (Saito et al., 2019; Saito et al., 2020; Xiu et al., 2022) to approximate surfaces, yet they often struggle with complex topologies and loose-fitting clothing. More recent works leverage advanced representations like 3DGS (Rong et al., 2025) and image-to-3D diffusion priors (Wang et al., 2025; Luo et al., 2024) to achieve impressive visual fidelity. However, these methods typically produce rigid meshes or unstructured point clouds that are inherently unsuitable for physical cloth simulation. While methods like Garment3DGen (Sarafianos et al., 2025) bridge this gap by utilizing template deformation to produce simulation-ready assets, they rely on a slow, computationally intensive per-asset optimization process. To obtain simulation-ready assets, a significant line of work attempts to infer 2D sewing patterns directly from 3D point clouds or images. NeuralTailor (Korosteleva and Lee, 2022) focused on reconstructing pattern structures from 3D geometry, and subsequent works (Liu et al., 2023; Chen et al., 2024) predict these patterns directly from single-view images using discriminative transformer architectures. To facilitate more structured generation, GarmentCode (Korosteleva and Sorkine-Hornung, 2023) introduced a programmatic domain-specific language (DSL) for sewing patterns, enabling parametric control. Building on this representation, subsequent research has increasingly focused on advanced generative modeling. DressCode (He et al., 2024) synthesizes novel patterns from text prompts using a GPT-based framework, while AIpparel (Nakayama et al., 2025) scales into a multimodal foundation model that handles complex pattern generation natively from mixed inputs. To further improve physical realism, Dress-1-to-3 (Li et al., 2025b) incorporates differentiable physics simulators to optimize geometric alignment of the inferred patterns. Concurrently, other state-of-the-art frameworks directly predict these programmatic structures from images leveraging large generative models (Bian et al., 2025; Zhou et al., 2025; Li et al., 2025a). While these methods have established sewing patterns as a robust domain for 3D animation, the synthesis of normalized, artifact-free textures for these panels remains an open challenge.

2.2. Generative 3D Texturing

The rise of large-scale vision-language models (Radford et al., 2021; Rombach et al., 2022; Saharia et al., 2022) has revolutionized 3D texture synthesis. Early optimization-based approaches either leverage CLIP (Radford et al., 2021) for optimizing texture maps of 3D models (Hong et al., 2022; Chen et al., 2022; Michel et al., 2022; Mohammad Khalid et al., 2022), or employ Score Distillation Sampling (SDS) to optimize texture maps by distilling gradients from a frozen 2D diffusion model (Poole et al., 2022; Metzer et al., 2023; Chen et al., 2023b). While capable of generating coherent global structures, these methods tend to hallucinate generic textures without high-frequency details. To improve fidelity, recent research has shifted toward projection-based inpainting, by utilizing depth-conditioned Stable Diffusion to progressively paint 3D meshes from multiple viewpoints (Richardson et al., 2023; Chen et al., 2023a; Cao et al., 2023). These frameworks iteratively project the current texture into screen space, inpaint missing regions using 2D priors, and project the result back. More recent systems like Paint3D (Zeng et al., 2024), MVPaint (Cheng et al., 2025) and Meta 3D TextureGen (Bensadoun et al., 2024) further refine this process by synchronizing multi-view generation or employing coarse-to-fine UV refinement to minimize seams. Despite their popularity, a fundamental limitation persists across these general-purpose methods: they rely on priors trained on natural photography. Consequently, they inherently entangle illumination with surface appearance, inevitably “baking in” transient lighting effects—such as cast shadows and specular highlights—directly into the UV map. In addition, the generated textures are subject to distortion and blurring due to the complex surface topology, leading to suboptimal results.

In the specific area of garment texturing, FabricDiffusion (Zhang et al., 2024) addresses some of these pitfalls by treating texture synthesis as a material extraction task. It utilizes a diffusion model to extract tileable, distortion-free material patches from a single image without lighting artifacts. However, because it relies on local texture tiling, FabricDiffusion fails to capture global structural information, such as specific graphic placements and non-uniform texture patterns across different UV islands. OmniFabric bridges this gap by shifting from local material extraction to global texture synthesis. We operate in the canonical UV space to disentangle base-color from illumination, synthesizing a holistic texture map that respects the global texture design in the reference image. While previous methods (Zeng et al., 2024; Bensadoun et al., 2024) also deploy a 2D refinement stage on the UV map, they only perform inpainting on occluded regions and fail to rectify distortion or remove baked-in artifacts, which are the domain specific challenges for garment texturing, as shown in Figure 2.

A figure showing the issues of existing 3D texturing methods.
Figure 2. Limits of baseline texturing methods. While existing methods can synthesize textures for 3D assets with decent visual quality, they fail to create garment textures free from distortions and baked-in illumination artifacts, rendering the asset unsuitable for downstream applications. A figure showing the issues of existing 3D texturing methods.

3. Method

A figure showing the overall pipeline of our method.
Figure 3. Overview of OmniFabric. Given a reference clothing image, OmniFabric synthesizes high-quality garment textures in sewing pattern UV space via a two-stage pipeline. Stage 1: We generate a coarse UV texture map by first predicting garment sewing patterns and reconstructing a canonical 3D garment. The input image texture is transferred onto the garment surface and expanded into multi-view using an off-the-shelf video generative model. These views are projected onto the canonical garment and reflected in UV space to produce a globally aligned coarse texture map. Stage 2: We refine the coarse texture with a DiT-based texture normalization network operating on UV patches. By conditioning on positional maps and spatially aligned tokens, the model removes baked-in artifacts and projection inconsistencies, producing clean and coherent sewing pattern textures that can be directly simulated into 3D garments.A figure showing the overall pipeline of our method.

Figure 3 presents an overview of our approach. We aim to reconstruct 3D garments with coherent textures from real-world images. We define the task and objective in Section 3.1. Next, we introduce a two-stage pipeline for generating holistic textures directly onto sewing patterns. This includes a coarse texture map initialization (Section 3.2) and a UV space texture normalization (Section 3.3). To train the model, we develop an automated pipeline to synthesize a large scale dataset of textured sewing patterns (Section 4).

3.1. Problem Definition

Given a single reference image II and an estimated 3D garment mesh ℳℛ\mathcal{M_{R}} parameterized by sewing patterns 𝒫\mathcal{P} from an off-the-shelf sewing pattern prediction model, our goal is to generate a normalized texture map 𝒯\mathcal{T} directly within the sewing pattern space. Unlike previous methods (Zhang et al., 2024) that focus on extracting local and repeatable material patches, we formulate our objective as a global texture synthesis task. We define 𝒯∈ℝH×W×3\mathcal{T}\in\mathbb{R}^{H\times W\times 3} as a representation of the normalized texture map of the garment, which is completely free of distortion, view-dependent illumination, cast shadows, and geometry-induced wrinkles. Following the paradigm of diffusion models, we formulate the synthesis of this holistic texture map as a conditional distribution mapping problem. We seek to learn a mapping function 𝒢\mathcal{G} such that:

(1) 𝒯∼𝒢⁡(ℐ,𝒫,ϵ),ϵ∼𝒩⁡(0,𝐈)\mathcal{T}\sim\mathcal{G}(\mathcal{I},\mathcal{P},\epsilon),\epsilon\sim\mathcal{N}(0,\mathbf{I})

We note that although 𝒯\mathcal{T} is not intrinsic albedo, it’s a close approximation and can be directly imported into cloth simulation engines, e.g., CLO (CLO Virtual Fashion Inc., 2024), with externally specified material assumption for the remaining properties including roughness, metallic and normal map, enabling the synthesis of simulation-ready 3D garments.

3.2. Coarse Texture Map Initialization

To guide texture synthesis in the sewing pattern space, we first aim to obtain a coarse texture initialization that fully leverages the visual information from the reference image. This requires establishing pixel-wise correspondences between the reference image and the sewing pattern. While this mapping is challenging, we recast the task via a VLM-driven and pose aware texture alignment step. We further leverage the multi-view generative priors of large video models to provide a complete and globally coherent texture initialization.

3.2.1. Pose aware texture alignment

Using Nano Banana Pro (Google, 2025), we repose the reference image II into I′I^{\prime} to perfectly align with the frontal silhouette of the rest-pose garment mesh ℳR\mathcal{M}_{R}. This alignment enables direct pixel projection of I′I^{\prime} onto ℳR\mathcal{M}_{R} and its corresponding sewing patterns 𝒫\mathcal{P}, producing the textured frontal render RfR_{f}. We observe that Nano Banana Pro successfully manipulates the input garment into a desired pose while preserving the original texture with high accuracy and fidelity.

A figure showing our data curation pipeline.
Figure 4. Framework of automated data creation. We first create a distorted texture image by applying a TPS warp to a sampled normalized texture. Both the normalized and distorted textures are then overlaid onto sampled sewing patterns. The textured sewing pattern can then be simulated into a training data point of rest-posed mesh ℳℛ\mathcal{M_{R}}, position map Pp​o​sP_{pos}, coarse texture map 𝒯′\mathcal{T}^{\prime} and frontal view rendering RfR_{f}. A figure showing our data curation pipeline.

3.2.2. Consistent multi-view synthesis.

After the initial pose alignment, a significant portion of the UV space occluded from the frontal view remains untextured. To hallucinate these missing regions while encouraging strong spatial continuity, we leverage a large-scale video generation model for multi-view generation. By conditioning the model on the textured frontal render RfR_{f}, we synthesize a 360-degree spinning video of the garment, naturally exploiting the model’s inherent temporal priors for structural consistency. Following recent projection-based texturing paradigms, we sample multi-view VV, four orthogonal views (front, back, left, and right) from the generated video. These observations are then projected onto ℳR\mathcal{M}_{R} and mapped to their corresponding UV coordinates on 𝒫\mathcal{P} to form the coarse texture map 𝒯′\mathcal{T}^{\prime}.

3.3. UV Space Texture Refinement

While 𝒯′\mathcal{T^{\prime}} preserves the coherent global texture of the garment, it contains baked-in artifacts like shadows, physical wrinkles and unpainted areas. Furthermore, spatial distortions in 𝒯′\mathcal{T^{\prime}} leave it far short of production-ready fabric quality, as shown in Figure 2. To address these issues, we introduce a texture normalization model designed to rectify 𝒯′\mathcal{T^{\prime}} into a normalized texture map 𝒯\mathcal{T}.

3.3.1. Training objective of multi-condition diffusion

We formulate texture normalization as a holistic UV synthesis problem by developing a conditional distribution mapping network. Unlike FabricDiffusion (Zhang et al., 2024), which relies on a single condition, i.e., a local textile patch, for the normalization task, our framework incorporates a multimodal conditioning set yy to guide global texture normalization. We extend Eq. 1 and minimize the following:

(2) ℒ=𝔼ℰ⁡(𝒯),y,ϵ∼𝒩⁡(0,𝐈),t​[‖ϵ−ϵθ​(zt,t,y)‖2],y={𝒯′,Rf,Pp​o​s}\mathcal{L}=\mathbb{E}_{\mathcal{E}(\mathcal{T}),y,\epsilon\sim\mathcal{N}(0,\mathbf{I}),t}\left[\|\epsilon-\epsilon_{\theta}(z_{t},t,y)\|^{2}\right],y=\{\mathcal{T}^{\prime},R_{f},P_{pos}\}

where ztz_{t} represents the noisy latent of the ground-truth texture at timestep tt. The multi-modal conditioning set yy includes the coarsely-textured map 𝒯′\mathcal{T^{\prime}}, textured frontal rendering RfR_{f} and a 3D position map Pp​o​sP_{pos}, created by projecting the vertices’ 3D positions of ℳℛ\mathcal{M_{R}} onto the 2D layout of 𝒫\mathcal{P}. Pp​o​sP_{pos} helps the model learn spatial connectivity between separate UV panels, ensuring seamless textures across fragmented UV islands.

3.3.2. Model architecture and training

To process multimodal conditioning set yy, we adopt a Diffusion Transformer (DiT) architecture following previous works (Tan et al., 2025a; Tan et al., 2025b), and leverage a unified token-processing strategy that treats all conditions as a unified sequence of tokens. We utilize a joint self-attention mechanism where the condition tokens (from 𝒯′\mathcal{T}^{\prime}, RfR_{f}, and Pp​o​sP_{pos}) and the noisy latent tokens from ztz_{t} interact within the same transformer blocks. Because 𝒯′\mathcal{T}^{\prime}, Pp​o​sP_{pos}, and the target 𝒯\mathcal{T} all share the same 2D panel layout, we apply a dynamic positional encoding strategy: image patches at the same spatial coordinates across these maps are assigned identical positional encodings, while only RfR_{f} retains its own coordinate system through distinct encodings. This minimal yet universal design allows the model to attend to the structural cues of Pp​o​sP_{pos} while simultaneously hallucinating missing textures in 𝒯′\mathcal{T}^{\prime}. We use a pretrained weight of DiT, and fine-tune the model on our synthetic textured sewing pattern data with LoRA (Hu et al., 2022).

4. Synthetic Training Data Creation

To train our texture normalization model, we develop a data engine (Figure 4) to curate a large-scale dataset of textured sewing patterns. The core objective is to synthesize training pairs that exhibit the complex appearance of real-world garments while providing ground-truth normalized texture maps for supervised learning.

4.1. Textured Sewing Pattern Creation

A key bottleneck in garment dataset construction is the lack of complex and high-resolution textures; most available texture sources provide only simple, repetitive patterns. To overcome this, we introduce a pipeline to produce diverse, high-fidelity texture images by using fashion images from real-world datasets, e.g., DeepFashion (Liu et al., 2016), as references. For each sample, we apply a Large Language Model (LLM) (Bai et al., 2023) to generate a detailed description encompassing texture layout, semantic elements, and color palettes. These descriptions, combined with the reference image, are fed into a vision-language model (Google, 2025) to synthesize textures with diverse designs. We then overlay the generated texture image onto sewing patterns sampled from GarmentCodeData. Note that the goal here is not to reconstruct the exact texture from the example image, but to generate plausible and realistic textures that provide sufficient diversity for training the texture synthesis model.

4.2. Training Pairs Construction

For each textured sewing pattern, we generate a training tuple (𝒯′,Rf,Pp​o​s,𝒯)(\mathcal{T}^{\prime},R_{f},P_{pos},\mathcal{T}) consisting of a coarse texture map, a frontal rendering, a positional map, and the ground-truth texture. To explicitly train the network to rectify geometric distortions, we apply a random Thin Plate Spline (TPS) (Bookstein, 1989) distortion to the initial texture image. The original, undistorted texture forms the ground-truth map 𝒯\mathcal{T}, while the TPS-distorted texture is overlaid onto the sewing pattern for physical simulation. Specifically, we perturb each control point pip_{i} of a regular grid fitted to the texture’s aspect ratio, giving p^i=pi+δi\hat{p}_{i}=p_{i}+\delta_{i} with

(3) δi=εi⋅λ⋅D800,εi∼𝒰⁡(−0.5,0.5)\delta_{i}=\varepsilon_{i}\cdot\lambda\cdot\frac{D}{800},\varepsilon_{i}\sim\mathcal{U}(-0.5,0.5)

where λ=25\lambda=25 by default controls the overall distortion magnitude and DD is the resolution of a texture image. A TPS is then fit between the regular grid {pi}\{p_{i}\} and its perturbed counterpart {p^i}\{\hat{p}_{i}\} and applied densely to warp the full-resolution texture. We stitch and drape the sewing pattern onto an A-pose SMPL body to yield a 3D garment mesh ℳR\mathcal{M}_{R}, from which we capture the frontal rendering RfR_{f}. The positional map Pp​o​sP_{pos} is then derived by projecting the x​y​zxyz coordinates of the mesh vertices onto the 2D panel layout 𝒫\mathcal{P}. Finally, we synthesize the coarse texture map 𝒯′\mathcal{T}^{\prime} by projecting RfR_{f} and a randomly sampled view RaR_{a} onto a neutral-posed mesh and unwrapping them back into the 2D pattern space. This explicitly bakes physics-induced artifacts and deformations into the initialization, compelling the model to learn undistortion and normalized base-color recovery.

5. Experiments

In this section, we present a comprehensive evaluation of OmniFabric through quantitative and qualitative analyses. We first detail the experimental setup, including curated synthetic dataset and metrics used for evaluation. We then demonstrate the effectiveness of our framework by comparing against SOTA methods for 3D texturing. Finally, we conduct ablation studies to validate our core architectural designs. Additional implementation details, extended qualitative results, our method’s adaptability to other sewing pattern prediction methods besides ChatGarment (Bian et al., 2025), examples of our curated garment data, and discussion on potential future directions are provided in the supplementary material.

5.1. Setup

5.1.1. Datasets

To train and evaluate our model, we construct a large-scale synthetic dataset of textured sewing patterns, as described in Section 4. We sample a total of 3K unique and diverse garment samples from GarmentCodeData (Korosteleva et al., 2024). We then use the pipeline in Figure 4 to generate texture images, and print them onto sampled sewing patterns, finally constructing a dataset of 30K textured sewing patterns with diverse appearances and structural designs (10 texture variations per garment style on average). We simulate garments using NVIDIA Warp (Macklin, 2022) with a per-sample random seed, so repeated simulations of the same sewing pattern produce distinct ℳR\mathcal{M}_{R} meshes and, consequently, different baked-in wrinkles in 𝒯′\mathcal{T}^{\prime}. For lighting, we sample random subsets of white point lights to introduce varied illumination and baked-in shading across the dataset. We train our model with this synthetic dataset, and keep a held-out set for quantitative comparison with other baseline models with a test/train ratio of 0.1. Besides synthetic data, we evaluate all methods on in-the-wild images, which is the major focus of this work. These in-the-wild images are obtained from DeepFashion (Liu et al., 2016) dataset and recent clothing-design images collected from the web. Due to the lack of ground truth in-the-wild images, we evaluate mainly through qualitative comparisons.

5.1.2. Metrics

We evaluate fidelity and structural coherence of the generated textures with a suite of standard metrics. We use LPIPS (Zhang et al., 2018) and DISTS (Ding et al., 2020) to measure visual similarity and structural consistency. We also compare with SSIM (Wang et al., 2004) and MS-SSIM (Wang et al., 2003), metrics employed to assess pixel-level structural accuracy. CLIP-score (CLIP-s) (Gal et al., 2022) measures the semantic alignment between the final textured garment and the reference images.

5.1.3. Baseline methods

We compare against several SOTA methods in 3D texturing and garment-specific texture transfer. Among them, Hunyuan3D-2.0 (Zhao et al., 2025) and Paint3D (Zeng et al., 2024) use multi-view generation to provide additional prior beyond the given single-view observation. FabricDiffusion (Zhang et al., 2024) is capable of creating fabric textures with no baked-in artifacts by normalizing local textile patterns with a diffusion model.

5.1.4. Implementation details

A figure showing the qualitative results.
Figure 5. Qualitative results of OmniFabric. Our method synthesizes 3D garments with high appearance consistency by predicting sewing patterns of normalized textures. In particular, OmniFabric handles intrinsic geometric distortions and reconstructs diverse texture types, including logos and printed graphics (1st and 2nd rows), and complex high-frequency patterns with dense visual details. Please zoom in to check the details. A figure showing the qualitative results.
A figure showing the qualitative comparisons among baselines.
Figure 6. Qualitative comparisons to state-of-the-art methods. OmniFabric is capable of generating accurate asymmetric textures and logos, while existing methods fail to preserve the high-frequency details on input cases with complex patterns. A figure showing the qualitative comparisons among baselines.

Our texture normalization model is based on the Diffusion Transformer (DiT) architecture, and we follow the multi-conditioning strategy of OminiControl (Tan et al., 2025a) by treating tokens from all input conditions as a unified token sequence. The model is fine-tuned using LoRA on a single NVIDIA A6000 GPU with a batch size of 8, keeping the VAE and primary DiT backbone fixed to preserve generative stability. We resize and pad sewing pattern maps to a resolution of 1024×\times1024 in a batch for training efficiency. During inference, we reverse the padding and resizing process, and deploy an image super-resolution model (Wang et al., 2021) to upscale the output. Since we are fine-tuning the model with LoRA, a higher training resolution can be easily achieved without GPU memory issues, making our framework scalable for training data of higher resolution. We fine-tune the model until convergence on our synthetic dataset. We utilize Veo 3 (Google DeepMind, 2025) as the multi-view video generation prior and Nano Banana Pro (Google, 2025) for the initial texture transfer to the frontal rendering RfR_{f}. We used ChatGarment (Bian et al., 2025) as the model for sewing pattern prediction throughout all experiments in the main paper. The average runtime for a single inference is roughly 5 minutes, depending on the current load on the Gemini model servers. The failure rate of Gemini models, e.g., generating unrelated textures on RfR_{f} or multi-views, is lower than 2%, estimated from randomly sampled generations. No additional filtering is necessary thanks to its stable performance. Additional details including prompts used for generation are provided in supplementary material.

5.2. Qualitative Comparisons

5.2.1. Texture synthesis by OmniFabric from in-the-wild images

We first show our results in Figure 5. Given a reference image, OmniFabric first generates a coarse texture map, serving as a roughly textured sewing pattern. Then, the texture normalization model removes the distortions and baked-in artifacts. From the front views of simulated 3D garments provided in Figure 5, we show that OmniFabric is capable of creating seamless textures across the 3D surface of garments. Note that while OmniFabric excels in textured garment synthesis from real-world images, we fine-tune our model only with synthetic data, which poses great scalability for model improvement.

5.2.2. Simulation results

We present simulations of synthesized 3D garments in Figure 9 via CLO. The sewing patterns generated by OmniFabric provide the normalized RGB base-color as albedo. Other parameters not predicted by our model—roughness, metallic, reflection intensity, and auto-generated normal map—are left at CLO’s default Fabric_Matte preset values. The Environment/HDRI maps for relighting examples are also chosen from CLO’s lighting presets. The simulated assets exhibit a convincing and photorealistic dynamic appearance, maintaining structural integrity and texture coherence even under extreme dynamics and varying illumination.

5.2.3. Comparison with state-of-the-art methods

We compare with other methods in Figure 6 using in-the-wild images as well. For input with complex textures and elements, both Hunyuan3D and Paint3D fail to preserve the high-frequency details. Even though FabricDiffusion can create normalized textile pattern and texture the given mesh by tiling local textile patches, it is limited to local textile patches and fails to generate coherent global appearance. Conversely, OmniFabric creates normalized textures while preserving coherent details.

5.3. Quantitative Comparisons

5.3.1. Synthetic dataset

We show quantitative results where all baselines are evaluated with our curated synthetic dataset for fair comparison. For each test case, we provide every baseline with the identical frontal rendering RfR_{f} and rest-posed mesh ℳℛ\mathcal{M_{R}}. Each method then performs its respective 3D texturing task on ℳℛ\mathcal{M_{R}}. To compute image-based metrics, the resulting 3D textures are projected back into the sewing pattern UV space to generate a comparable texture map 𝒯\mathcal{T} for each baseline. Table 1 shows OmniFabric excels in all metrics, particularly in preserving global structural coherence and eliminating projection artifacts.

5.3.2. Real-world data

Table 1. Quantitative comparisons. We compare with baselines on our synthetic dataset. The results show that OmniFabric greatly surpasses all other methods in terms of both pixel and perceptual quality.
LPIPS ↓\downarrow SSIM ↑\uparrow MS-SSIM ↑\uparrow DISTS ↓\downarrow CLIP-s ↑\uparrow
FabricDiffusion (Zhang et al., 2024) 0.273 0.645 0.707 0.259 0.906
Paint3D (Zeng et al., 2024) 0.311 0.657 0.704 0.266 0.890
Hunyuan3D-2.0 (Zhao et al., 2025) 0.223 0.712 0.717 0.220 0.924
OmniFabric (ours) 0.092 0.868 0.905 0.121 0.963
Table 2. User study for real-world data. We compare with baselines on real-world input images in 4 aspects: overall quality, fidelity, back-view plausibility and global coherence.
Overall rank ↓\downarrow Fidelity ↑\uparrow Back-view ↑\uparrow Global Coherence ↑\uparrow
FabricDiffusion (Zhang et al., 2024) 3.40 1.75 2.69 3.15
Paint3D (Zeng et al., 2024) 3.41 1.86 2.81 3.03
Hunyuan3D-2.0 (Zhao et al., 2025) 2.10 3.43 3.41 3.50
OmniFabric (ours) 1.09 4.44 4.35 4.41

Besides comparing with other methods in Figure 6, we conduct a user study to evaluate the performances of the methods due to the absence of real-world dataset of textured sewing patterns. As detailed in Table 2, the study involves 13 participants evaluating results from 10 real-world image inputs based on four criteria: overall quality ranking (1 to 4, lower is better), fidelity to the input appearance, back-view plausibility, and global texture coherence (1 to 5, higher is better). OmniFabric achieves the best results across all criteria, demonstrating that it is preferred on real images and more faithfully preserves coherent garment textures.

5.4. Reproducibility via Open-Source Models

Table 3. We construct a fully open-source alternative and it achieves similar performance to the default settings using closed-source models, showing the great reproducibility of our framework.
LPIPS ↓\downarrow SSIM ↑\uparrow DISTS ↓\downarrow CLIP-s ↑\uparrow
OmniFabric w/ default settings 0.092 0.868 0.121 0.963
OmniFabric w/ open-source models 0.114 0.847 0.149 0.955
Table 4. We analyze error accumulation across certain modules by replacing intermediate outputs in our model pipeline with ground-truth (GT) counterparts. “Pred.” and “GT” mean the intermediate output is provided by Gemini models and GT respectively, and "SP Norm." refers to sewing pattern normalization. Results show that our texture normalization model effectively removes the accumulated error.
Replacement Setting LPIPS ↓\downarrow SSIM ↑\uparrow DISTS ↓\downarrow CLIP-s ↑\uparrow
   (a) GT I′I^{\prime} + GT VV 0.105 0.849 0.159 0.944
   (b) Pred. I′I^{\prime} + GT VV 0.121 0.837 0.176 0.935
   (c) Pred. I′I^{\prime} + Pred. VV 0.144 0.804 0.166 0.938
   (d) Pred. I′I^{\prime} + Pred. VV + SP Norm. 0.092 0.868 0.121 0.963
Table 5. Ablation study shows that our designs are essential for the texture normalization model to generate position-aligned details with high fidelity.
Setting LPIPS ↓\downarrow SSIM ↑\uparrow DISTS ↓\downarrow CLIP-s ↑\uparrow
OmniFabric (ours) 0.092 0.868 0.121 0.963
   (a) w/o coarse texture map 𝒯′\mathcal{T^{\prime}} 0.407 0.601 0.306 0.887
   (b) w/o frontal view rendering RfR_{f} 0.097 0.850 0.129 0.962
   (c) w/o position map Pp​o​sP_{pos} 0.098 0.848 0.130 0.961

While using Gemini models for pose transfer and multi-view generation (Section 3.2) as default implementation of OmniFabric, we further construct a fully open-source alternative to facilitate reproducibility. To replace the closed-source models, we use FLUX.2 (Labs, 2025) as image generation backbone and incorporate Ministral 3 (Liu et al., 2026) for prompt up-sampling. We feed textured frontal view RfR_{f} and silhouettes of four orthogonal views as conditions to generate the multi-view with this alternative. Table 3 shows that open-source alternative achieves similar performance with the same experiment settings in Section 5.3.

5.5. Error-Accumulation Analysis

We analyze error accumulation across modules in our model architecture in this section by progressively replacing intermediate outputs, specifically the reposed reference image I′I^{\prime} and generated multi-view VV, with ground-truth counterparts. As shown in Table 4, settings (a) to (c) accumulate errors across stages including texture projection, pose alignment and multi-view generation, and (d) shows that our normalization model effectively fixes these aggregated artifacts on sewing patterns.

5.6. Ablation Studies and Analyses

5.6.1. Importance of the canonical 3D and coarse texture map

As shown in Table 5-(a), without the coarse texture map, which serves as the distillation of generative priors, the performance drops significantly. This shows that the projected textures do provide essential guidance for the texture generation. Figure 8 shows that the model is randomly generating textures that “look similar” instead of pixel-aligned content without coarse texture map 𝒯′\mathcal{T^{\prime}}.

5.6.2. Effect of the frontal view rendering condition

We observe that removing RfR_{f} during training leads to inferior results (Table 5-(b)), suggesting that the frontal-view rendering provides important cues for the model. In particular, RfR_{f} offers a holistic visual reference of the garment appearance, helping the model better understand global layout and consistency of the texture across sewing pattern pieces.

5.6.3. Effect of position map

As shown in Table 5-(c), without Pp​o​sP_{pos}, the model not only shows suboptimal results but also textures and elements in wrong positions. As illustrated in Figure 8, a lemon is mistakenly generated in a blank region of the sewing pattern.

5.6.4. Importance of geometric distortion in dataset.

The right side of Figure 8 shows the necessity of creating geometric distortion in the synthetic training data. Through training to remove geometric distortion, the model can not only generate normalized structured patterns on the sleeves, e.g., the checker pattern, but also produce clean appearances for other elements, e.g., texts and logos.

5.7. Limitations

A figure showing the failure cases of our method.
Figure 7. Limitations. OmniFabric may synthesize mismatched textures from inaccurately generated multi-views (1st row), or generate garments with incorrect geometry even though appearance is well-transferred (2nd row).A figure showing the failure cases of our method.

As shown in Figure 7, correctness of synthesized textures can be affected by inaccurately generated multi-views (1st row). OmniFabric is also unable to refine the mismatched geometry ℳR\mathcal{M}_{R} predicted by ChatGarment (Bian et al., 2025) (2nd row); however, it does not change the role of our refinement model: given a sewing pattern, it normalizes the textures by reducing distortions, wrinkles, shadows, and projection artifacts, regardless of specific garment geometry. We put more discussion of limitations in the supplementary material.

6. Conclusion and Discussion

We introduce OmniFabric, a novel framework for synthesizing 3D garments with normalized and globally coherent textures. By leveraging large-scale generative priors and training a texture normalization model, we effectively disentangle intrinsic albedo from view-dependent artifacts including baked-in shadows and wrinkles. Our method bridges the gap between single-view observations and simulation-ready 3D assets, maintaining structural integrity across complex sewing patterns. For this study, we focus on the synthesis of sewing patterns with normalized textures within the GarmentCode framework. An immediate and promising extension would be to consider the joint prediction of complete PBR material maps, including roughness, metallic, and normal maps, alongside the albedo. We believe this direction will further enhance the photorealism of the reconstructed garments under diverse environmental lighting, providing even greater utility for immersive digital content creation.

A figure showing the ablation study of our method.
Figure 8. Qualitative analysis of key design mechanisms. We demonstrate the effectiveness of the core designs of our texture normalization model. Without 𝒯′\mathcal{T^{\prime}}, the model fails to generate consistent appearance but arbitrary textures with similar colors and elements. We also observe that RfR_{f} and Pp​o​sP_{pos} are both essential conditions for the model to generate pixel-aligned content observed from the input image (left figure). We also show that our method can effectively rectify distortion when trained with distortion-augmented dataset (right figure). A figure showing the ablation study of our method.
A figure showing the simulation results of our synthesized 3D garment.
Figure 9. Simulation results of our generated 3D garments. 3D garments generated by our method can be directly simulated under varying human pose and lighting conditions and show convincing results. A figure showing the simulation results of our synthesized 3D garment.

References

  • Bai et al. (2023) J. Bai, S. Bai, Y. Chu, Z. Cui, K. Dang, X. Deng, Y. Fan, W. Ge, Y. Han, F. Huang, et al. Qwen technical report. arXiv preprint arXiv:2309.16609. Cited by: §4.1.
  • Bensadoun et al. (2024) R. Bensadoun, Y. Kleiman, I. Azuri, O. Harosh, A. Vedaldi, N. Neverova, and O. Gafni Meta 3d texturegen: fast and consistent texture generation for 3d objects. arXiv preprint arXiv:2407.02430. Cited by: §1, §2.2, §2.2.
  • Bian et al. (2025) S. Bian, C. Xu, Y. Xiu, A. Grigorev, Z. Liu, C. Lu, M. J. Black, and Y. Feng Chatgarment: garment estimation, generation and editing via large language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 2924–2934. Cited by: §1, §2.1, §5.1.4, §5.7, §5.
  • Bookstein (1989) F. L. Bookstein Principal warps: thin-plate splines and the decomposition of deformations. IEEE Transactions on pattern analysis and machine intelligence 11 (6), pp. 567–585. Cited by: §4.2.
  • Cao et al. (2023) T. Cao, K. Kreis, S. Fidler, N. Sharp, and K. Yin Texfusion: synthesizing 3d textures with text-guided image diffusion models. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 4169–4181. Cited by: §2.2.
  • Chen et al. (2024) C. Chen, J. Su, M. Hu, C. Yao, and H. Chu Panelformer: sewing pattern reconstruction from 2d garment images. In Proceedings of the IEEE/CVF winter conference on applications of computer vision, pp. 454–463. Cited by: §2.1.
  • Chen et al. (2023a) D. Z. Chen, Y. Siddiqui, H. Lee, S. Tulyakov, and M. Nießner Text2tex: text-driven texture synthesis via diffusion models. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 18558–18568. Cited by: §1, §2.2.
  • Chen et al. (2023b) R. Chen, Y. Chen, N. Jiao, and K. Jia Fantasia3d: disentangling geometry and appearance for high-quality text-to-3d content creation. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 22246–22256. Cited by: §2.2.
  • Chen et al. (2022) Y. Chen, R. Chen, J. Lei, Y. Zhang, and K. Jia Tango: text-driven photorealistic and robust 3d stylization via lighting decomposition. Advances in neural information processing systems 35, pp. 30923–30936. Cited by: §2.2.
  • Cheng et al. (2025) W. Cheng, J. Mu, X. Zeng, X. Chen, A. Pang, C. Zhang, Z. Wang, B. Fu, G. Yu, Z. Liu, et al. Mvpaint: synchronized multi-view diffusion for painting anything 3d. In Proceedings of the Computer Vision and Pattern Recognition Conference, pp. 585–594. Cited by: §2.2.
  • CLO Virtual Fashion Inc. (2024) CLO Virtual Fashion Inc. CLO. Note: https://www.clo3d.comVersion 7.2 Cited by: §3.1.
  • Ding et al. (2020) K. Ding, K. Ma, S. Wang, and E. P. Simoncelli Image quality assessment: unifying structure and texture similarity. IEEE transactions on pattern analysis and machine intelligence 44 (5), pp. 2567–2581. Cited by: §5.1.2.
  • Gal et al. (2022) R. Gal, Y. Alaluf, Y. Atzmon, O. Patashnik, A. H. Bermano, G. Chechik, and D. Cohen-Or An image is worth one word: personalizing text-to-image generation using textual inversion. arXiv preprint arXiv:2208.01618. Cited by: §5.1.2.
  • Google DeepMind (2025) Google DeepMind Veo: google’s most capable video generation model. Note: https://deepmind.google/models/veo/Accessed: September 24, 2026 Cited by: §1, §5.1.4.
  • Google (2025) Google Gemini 3 pro image. Note: https://gemini.google.com/ Cited by: §1, §3.2.1, §4.1, §5.1.4.
  • Habermann et al. (2021) M. Habermann, L. Liu, W. Xu, M. Zollhoefer, G. Pons-Moll, and C. Theobalt Real-time deep dynamic characters. ACM Transactions on Graphics (ToG) 40 (4), pp. 1–16. Cited by: §1.
  • He et al. (2024) K. He, K. Yao, Q. Zhang, J. Yu, L. Liu, and L. Xu Dresscode: autoregressively sewing and generating garments from text guidance. ACM Transactions on Graphics (TOG) 43 (4), pp. 1–13. Cited by: §1, §2.1.
  • Hong et al. (2022) F. Hong, M. Zhang, L. Pan, Z. Cai, L. Yang, and Z. Liu Avatarclip: zero-shot text-driven generation and animation of 3d avatars. arXiv preprint arXiv:2205.08535. Cited by: §2.2.
  • Hu et al. (2022) E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen, et al. Lora: low-rank adaptation of large language models.. Iclr 1 (2), pp. 3. Cited by: §3.3.2.
  • Hughes et al. (2007) C. J. Hughes, R. Grzeszczuk, E. Sifakis, D. Kim, S. Kumar, A. P. Selle, J. Chhugani, M. Holliman, and Y. Chen Physical simulation for animation and visual effects: parallelization and characterization for chip multiprocessors. ACM SIGARCH Computer Architecture News 35 (2), pp. 220–231. Cited by: §1.
  • Hwangbo et al. (2020) H. Hwangbo, E. H. Kim, S. Lee, and Y. J. Jang Effects of 3d virtual “try-on” on online sales and customers’ purchasing experiences. Ieee Access 8, pp. 189479–189489. Cited by: §1.
  • Kim and LaBat (2013) D. Kim and K. LaBat Consumer experience in using 3d virtual garment simulation technology. Journal of the Textile Institute 104 (8), pp. 819–829. Cited by: §1.
  • Korosteleva et al. (2024) M. Korosteleva, T. L. Kesdogan, F. Kemper, S. Wenninger, J. Koller, Y. Zhang, M. Botsch, and O. Sorkine-Hornung GarmentCodeData: a dataset of 3d made-to-measure garments with sewing patterns. In European Conference on Computer Vision, pp. 110–127. Cited by: §5.1.1.
  • Korosteleva and Lee (2022) M. Korosteleva and S. Lee Neuraltailor: reconstructing sewing pattern structures from 3d point clouds of garments. ACM Transactions on Graphics (TOG) 41 (4), pp. 1–16. Cited by: §2.1.
  • Korosteleva and Sorkine-Hornung (2023) M. Korosteleva and O. Sorkine-Hornung Garmentcode: programming parametric sewing patterns. ACM Transactions on Graphics (TOG) 42 (6), pp. 1–15. Cited by: §2.1.
  • Labs (2025) B. F. Labs FLUX.2: Frontier Visual Intelligence. Note: https://bfl.ai/blog/flux-2 Cited by: §5.4.
  • Li et al. (2025a) X. Li, Q. Yao, and Y. Wang GarmentDiffusion: 3d garment sewing pattern generation with multimodal diffusion transformers. arXiv preprint arXiv:2504.21476. Cited by: §2.1.
  • Li et al. (2025b) X. Li, C. Yu, W. Du, Y. Jiang, T. Xie, Y. Chen, Y. Yang, and C. Jiang Dress-1-to-3: single image to simulation-ready 3d outfit with diffusion prior and differentiable physics. ACM Transactions on Graphics (TOG) 44 (4), pp. 1–16. Cited by: §1, §2.1.
  • Liu et al. (2026) A. H. Liu, K. Khandelwal, S. Subramanian, V. Jouault, A. Rastogi, A. Sadé, A. Jeffares, A. Jiang, A. Cahill, A. Gavaudan, et al. Ministral 3. arXiv preprint arXiv:2601.08584. Cited by: §5.4.
  • Liu et al. (2023) L. Liu, X. Xu, Z. Lin, J. Liang, and S. Yan Towards garment sewing pattern reconstruction from a single image. ACM Transactions on Graphics (TOG) 42 (6), pp. 1–15. Cited by: §2.1.
  • Liu et al. (2016) Z. Liu, P. Luo, S. Qiu, X. Wang, and X. Tang DeepFashion: powering robust clothes recognition and retrieval with rich annotations. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: §4.1, §5.1.1.
  • Luo et al. (2024) Z. Luo, H. Liu, C. Li, W. Du, Z. Jin, W. Sun, Y. Nie, W. Chen, and X. Han Garverselod: high-fidelity 3d garment reconstruction from a single in-the-wild image using a dataset with levels of details. ACM Transactions on Graphics (TOG) 43 (6), pp. 1–12. Cited by: §2.1.
  • Macklin (2022) M. Macklin Warp: A High-performance Python Framework for GPU Simulation and Graphics. Note: NVIDIA GPU Technology Conference (GTC) External Links: Link Cited by: §5.1.1.
  • Metzer et al. (2023) G. Metzer, E. Richardson, O. Patashnik, R. Giryes, and D. Cohen-Or Latent-nerf for shape-guided generation of 3d shapes and textures. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 12663–12673. Cited by: §2.2.
  • Michel et al. (2022) O. Michel, R. Bar-On, R. Liu, S. Benaim, and R. Hanocka Text2mesh: text-driven neural stylization for meshes. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 13492–13502. Cited by: §2.2.
  • Mohammad Khalid et al. (2022) N. Mohammad Khalid, T. Xie, E. Belilovsky, and T. Popa Clip-mesh: generating textured meshes from text using pretrained image-text models. In SIGGRAPH Asia 2022 conference papers, pp. 1–8. Cited by: §2.2.
  • Nakayama et al. (2025) K. Nakayama, J. Ackermann, T. L. Kesdogan, Y. Zheng, M. Korosteleva, O. Sorkine-Hornung, L. J. Guibas, G. Yang, and G. Wetzstein Aipparel: a multimodal foundation model for digital garments. In Proceedings of the Computer Vision and Pattern Recognition Conference, pp. 8138–8149. Cited by: §C.2, §1, §2.1.
  • Poole et al. (2022) B. Poole, A. Jain, J. T. Barron, and B. Mildenhall Dreamfusion: text-to-3d using 2d diffusion. arXiv preprint arXiv:2209.14988. Cited by: §2.2.
  • Radford et al. (2021) A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pp. 8748–8763. Cited by: §2.2.
  • Richardson et al. (2023) E. Richardson, G. Metzer, Y. Alaluf, R. Giryes, and D. Cohen-Or Texture: text-guided texturing of 3d shapes. In ACM SIGGRAPH 2023 conference proceedings, pp. 1–11. Cited by: §1, §2.2.
  • Rombach et al. (2022) R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 10684–10695. Cited by: §2.2.
  • Rong et al. (2025) B. Rong, A. Grigorev, W. Wang, M. J. Black, B. Thomaszewski, C. Tsalicoglou, and O. Hilliges Gaussian garments: reconstructing simulation-ready clothing with photorealistic appearance from multi-view video. In 2025 International Conference on 3D Vision (3DV), pp. 1054–1063. Cited by: §1, §2.1.
  • Saharia et al. (2022) C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Denton, K. Ghasemipour, R. Gontijo Lopes, B. Karagol Ayan, T. Salimans, et al. Photorealistic text-to-image diffusion models with deep language understanding. Advances in neural information processing systems 35, pp. 36479–36494. Cited by: §2.2.
  • Saito et al. (2019) S. Saito, Z. Huang, R. Natsume, S. Morishima, A. Kanazawa, and H. Li Pifu: pixel-aligned implicit function for high-resolution clothed human digitization. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 2304–2314. Cited by: §2.1.
  • Saito et al. (2020) S. Saito, T. Simon, J. Saragih, and H. Joo Pifuhd: multi-level pixel-aligned implicit function for high-resolution 3d human digitization. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 84–93. Cited by: §2.1.
  • Sarafianos et al. (2025) N. Sarafianos, T. Stuyck, X. Xiang, Y. Li, J. Popovic, and R. Ranjan Garment3dgen: 3d garment stylization and texture generation. In 2025 International Conference on 3D Vision (3DV), pp. 1382–1393. Cited by: §1, §2.1.
  • Tan et al. (2025a) Z. Tan, S. Liu, X. Yang, Q. Xue, and X. Wang Ominicontrol: minimal and universal control for diffusion transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 14940–14950. Cited by: §3.3.2, §5.1.4.
  • Tan et al. (2025b) Z. Tan, Q. Xue, X. Yang, S. Liu, and X. Wang Ominicontrol2: efficient conditioning for diffusion transformers. arXiv preprint arXiv:2503.08280. Cited by: §3.3.2.
  • Wang et al. (2021) X. Wang, L. Xie, C. Dong, and Y. Shan Real-esrgan: training real-world blind super-resolution with pure synthetic data. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 1905–1914. Cited by: §5.1.4.
  • Wang et al. (2025) Y. Wang, C. Zhang, G. Frazão, J. Yang, A. Ichim, T. Beeler, and F. De la Torre GarmentCrafter: progressive novel view synthesis for single-view 3d garment reconstruction and editing. arXiv preprint arXiv:2503.08678. Cited by: §1, §2.1.
  • Wang et al. (2004) Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing 13 (4), pp. 600–612. Cited by: §5.1.2.
  • Wang et al. (2003) Z. Wang, E. P. Simoncelli, and A. C. Bovik Multiscale structural similarity for image quality assessment. In The Thirty-Seventh Asilomar Conference on signals, systems & computers, 2003, Vol. 2, pp. 1398–1402. Cited by: §5.1.2.
  • Xiu et al. (2022) Y. Xiu, J. Yang, D. Tzionas, and M. J. Black Icon: implicit clothed humans obtained from normals. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 13286–13296. Cited by: §2.1.
  • Zeng et al. (2024) X. Zeng, X. Chen, Z. Qi, W. Liu, Z. Zhao, Z. Wang, B. Fu, Y. Liu, and G. Yu Paint3d: paint anything 3d with lighting-less texture diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 4252–4262. Cited by: §1, §2.2, §2.2, §5.1.3, Table 1, Table 2.
  • Zhang et al. (2024) C. Zhang, Y. Wang, F. Vicente, C. Wu, J. Yang, T. Beeler, and F. De la Torre FabricDiffusion: high-fidelity texture transfer for 3d garments generation from in-the-wild images. In SIGGRAPH Asia 2024 Conference Papers, pp. 1–12. Cited by: §1, §2.2, §3.1, §3.3.1, §5.1.3, Table 1, Table 2.
  • Zhang et al. (2018) R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 586–595. Cited by: §5.1.2.
  • Zhao et al. (2025) Z. Zhao, Z. Lai, Q. Lin, Y. Zhao, H. Liu, S. Yang, Y. Feng, M. Yang, S. Zhang, X. Yang, et al. Hunyuan3d 2.0: scaling diffusion models for high resolution textured 3d assets generation. arXiv preprint arXiv:2501.12202. Cited by: §1, §5.1.3, Table 1, Table 2.
  • Zhou et al. (2025) F. Zhou, R. Liu, C. Liu, G. He, Y. Li, X. Jin, and H. Wang Design2GarmentCode: turning design concepts to tangible garments through program synthesis. In Proceedings of the Computer Vision and Pattern Recognition Conference, pp. 23712–23722. Cited by: §2.1.

Supplementary Material

Supplementary Material A A Details of Dataset Construction

A primary contribution of this work is the automated data engine designed to generate high-fidelity textured sewing patterns. While high-detail manual texturing is possible, the process remains prohibitively time-consuming for large-scale applications. In contrast, our pipeline efficiently produces diverse garment styles and textures with a vast range of complexity, facilitating the construction of digital garment datasets at scale. This automated engine is a core contribution that addresses the scarcity of complex, high-resolution textures in existing garment research. We have included several representative examples in this supplement and will release the full dataset as well as the curation pipeline.

A figure showing several examples in our curated dataset.
Figure S1. Data examples of our synthetic dataset. We show the synthesized textured sewing pattern and its simulated 3D garment for each datapoint.A figure showing several examples in our curated dataset.

A.1. Reordering Sewing Pattern for Texturing

To align our generation process with real-world manufacturing, we introduce a reordering strategy for sewing pattern layouts. In professional garment construction, textures are often seamless across specific functional groups—such as a large frontal fabric panel—but exhibit discontinuities at structural seams, such as the transition from a sleeve to the chest. These discontinuities arise because 3D-stitched edges are often not geometrically matched in 2D space (e.g., a curved sleeve head joining a straight armhole), necessitating different fabric cuts. Following these observations, our pipeline first reorders the 2D sewing pattern pieces by merging related panels, such as the front and back pieces of a single sleeve, into unified spatial groups. By overlaying textures onto this reordered 2D layout, we ensure seamless pattern continuity within each functional group while maintaining the realistic texture breaks required for production-ready 3D assets.

Supplementary Material B B Additional Implementation Details

In this section, we provide further technical specifics regarding our texturing pipeline. As detailed in the main paper, we utilize a single-view garment observation to generate a 360-degree rotation video. From this sequence, we extract four keyframes corresponding to the orthogonal front, back, left, and right camera views for the multi-view texture projection process. Given that the video depicts a garment in a canonical A-pose, these four views are sufficient to capture the vast majority of surface details. The rest-posed mesh, ℳℛ\mathcal{M}_{\mathcal{R}}, is reconstructed by stitching and draping the predicted sewing patterns onto a A-posed SMPL human body model.

B.1. Non-overlapping Texture Projection

A key distinction between our projection method and state-of-the-art baselines, such as Hunyuan3D-2.0, lies in the source and application of multi-view data. While traditional methods project independent images generated by multi-view diffusion models, we project selected frames from a structurally consistent video prior using a specialized non-overlapping strategy. We observe that while multi-view diffusion models improve cross-view alignment, they often sacrifice high-frequency details to maintain that consistency. To preserve these details, our non-overlapping strategy begins by projecting textures from the front and back views—the two perspectives with minimal observational overlap—directly onto the mesh. Subsequently, textures from the side views are projected exclusively onto previously untextured regions, effectively performing a projection-based inpainting. Finally, our data-driven texture normalization model rectifies any minor artifacts or discontinuities at the seams of these projected textures. This approach allows OmniFabric to retain intricate, high-frequency surface details while ensuring global multi-view consistency.

B.2. Prompt Configurations

In this section we specify the design of prompts for the Gemini models in our framework. We use the prompt below for pose aware texture alignment, which generates a textured frontal view of rest-pose garment mesh ℳR\mathcal{M}_{R}:

The first apparel is my target apparel. Please generate textures on this apparel, so the textures are identical to the apparel in the reference image. Please don’t change the style, size and pose of the target apparel.

For the multi-view generation with video model, we use the following prompt:

Create a continuous 360 degree rotation video of this apparel, making it rotate horizontally 360 degree, like a microwave.

B.3. Training Details

We fine-tune all three baselines from pretrained weights, adapting the components described below. For Paint3D, we fine-tune its position encoder, as specified in its paper. For Hunyuan3D, we fine-tune its multi-view image generator. And for FabricDiffusion, we train its texture generator using paired textile data. All the models are fine-tuned upon pre-trained model weights.

Supplementary Material C C Additional Results and Analyses

A figure showing that our method is compatible with other sewing pattern prediction methods.
Figure S2. Adaptability to other sewing pattern prediction method. The simulation results show that OmniFabric can be seamlessly integrated with other methods of sewing pattern prediction. A figure showing that our method is compatible with other sewing pattern prediction methods.

C.1. Robustness to Structural Variations in Sewing Patterns

A figure showing qualitative results of our method.
Figure S3. Qualitative results. OmniFabric effectively removes distortion artifacts and geometry-induced wrinkles, generating normalized textured sewing patterns that are simulation-ready.A figure showing qualitative results of our method.

We present further qualitative evaluations in Figure S3, which demonstrate the robustness of our framework across a variety of garment styles. The results illustrate that our texture normalization model effectively rectifies the coarse texture map by eliminating projection-induced distortions and removing physics-based wrinkles inherent in the initial capture. While the geometry simulated from the predicted sewing patterns may occasionally deviate from the exact silhouette in the input image, we emphasize that such discrepancies arise from the limitations of the underlying sewing pattern prediction method rather than the texturing process. To mitigate this, our pipeline employs a texture transferring model that aligns the input image textures with the rendered silhouette of the predicted garment. This strategy successfully bridges the gap between the predicted geometry and the original observation, ensuring that the synthesized textures remain globally coherent and structurally aligned.

C.2. Compatibility with Other Sewing Pattern Prediction Method

While the main paper utilizes ChatGarment for sewing pattern prediction, our framework is designed to be agnostic to the specific reconstruction method employed. To demonstrate this flexibility, we evaluated our pipeline using sewing patterns predicted by Alpparel (Nakayama et al., 2025), a recent state-of-the-art approach based on Large Vision-Language Models. As illustrated in Figure S2, OmniFabric consistently generates textured garments with aligned geometry and high-frequency details, confirming that our synthesis capability is not restricted to a single pattern reconstruction architecture.It should be noted, however, that the simulated garment geometry may occasionally exhibit minor mismatches relative to the original input image due to imperfect pattern prediction. As discussed in Section C.1, our methodology specifically focuses on maintaining global texture consistency even when the underlying 3D mesh, which is simulated from predicted sewing pattern, is not perfectly aligned with the reference observation. This robustness ensures that OmniFabric can be seamlessly integrated with future advancements in sewing pattern prediction.

Supplementary Material D D Limitations and Future Work

While OmniFabric represents a significant advancement in garment texturing, several limitations remain that offer promising avenues for future research. First, although our model effectively removes transient lighting artifacts to produce a clean appearance, it does not yet explicitly disentangle albedo from shading in a strictly principled, physics-based manner, a distinction that should be noted to avoid overstating our current appearance decomposition capabilities. Furthermore, because the framework relies on generative priors, unseen or occluded regions are occasionally hallucinated; however, the system is not strictly limited to single-view observations and could incorporate multiple frames or video sequences in practice to reduce these hallucinations and improve reconstruction fidelity. Finally, certain materials with strong view-dependent effects, such as highly reflective leather or metallic fabrics, remain challenging to reconstruct faithfully, suggesting that future iterations should incorporate the joint prediction of complete PBR material maps—including roughness and normal maps—to achieve true photorealism across diverse environmental lighting conditions.