OmniFabric: Coherent UV Space Texture Synthesis for 3D Garment Reconstruction

Abstract.
Automated generation of production-ready 3D garment assets from a single image is a central challenge in digital content creation. While recent generative models have significantly advanced 3D geometry reconstruction, synthesizing high-quality textures remains a bottleneck. Existing methods often bake environmental illumination and shadows directly into the texture map, or they fail to maintain global structural coherence, making the resulting assets unusable for physical simulation and relighting. In this work, we introduce OmniFabric, a novel approach that synthesizes globally coherent texture maps directly within the 2D sewing pattern space. Given a single reference image, our pipeline utilizes an estimated 3D mesh and generative priors of powerful Vision-Language Models (VLM) to establish a complete but coarse texture initialization across the unwrapped sewing patterns. We then leverage a specialized diffusion transformer, trained via an automated synthetic data engine and conditioned on 3D positional features, to refine this initialization directly in the canonical UV domain. This effectively removes distortion and baked-in artifacts to extract a clean and normalized texture map that preserves the original garment design. Extensive experiments demonstrate that OmniFabric significantly outperforms state-of-the-art baselines, yielding photorealistic 3D garments with high-quality textures. Project page: https://humansensinglab.github.io/OmniFabric
Keywords:
Texture Synthesis, 3D Garment Reconstruction, Garment Sewing Patterns1. Introduction
Creating high-fidelity, simulation-ready 3D garments is essential for modern digital production, with applications spanning gaming (Habermann et al., 2021), e-commerce (Hwangbo et al., 2020; Kim and LaBat, 2013), and cinematic visual effects (Hughes et al., 2007). While traditional asset creation relies on labor-intensive manual modeling, there is growing interest in automating this process to synthesize assets directly from images. Despite recent strides in reconstructing accurate 3D garment geometry (Bian et al., 2025; Nakayama et al., 2025; Sarafianos et al., 2025; He et al., 2024; Li et al., 2025b; Rong et al., 2025; Wang et al., 2025), generating production-ready textures from a single observation remains an open problem.
Garment texture synthesis is uniquely difficult due to the complex dynamics of fabric, which frequently produce deep folds, self-occlusions, and extreme pose variations (see Figure 1). To be truly usable in downstream applications, a texture must not only look realistic in a static view but also be free from baked-in lighting and geometric distortion, capturing the garment’s intrinsic albedo. However, current state-of-the-art 3D texturing frameworks (Richardson et al., 2023; Chen et al., 2023a; Zeng et al., 2024; Zhao et al., 2025; Bensadoun et al., 2024) typically rely on multi-view diffusion priors to iteratively “paint” assets via direct projection onto draped 3D geometry. Consequently, unwrapping these folded surfaces inevitably causes large geometric distortion and occlusion. Furthermore, these models fail to disentangle intrinsic material properties from environmental illumination, baking transient lighting artifacts, such as shadows and geometry-induced wrinkles, directly into the texture map.
To circumvent these issues, recent work such as FabricDiffusion (Zhang et al., 2024) has explored texture synthesis directly within the regularized, flat 2D sewing pattern domain, successfully yielding simulation-ready texture maps. However, FabricDiffusion focuses exclusively on local texture synthesis by extracting tileable material patches from the source image, ignoring the global texture layout. Consequently, it struggles at reconstructing asymmetric designs, large-scale prints, or irregular spatial patterns. We argue that synthesizing a globally coherent texture map that faithfully reflects the visual identity of the input image requires leveraging the complete context of the reference observation, rather than relying on isolated patches. This presents a fundamental challenge:
How do we establish pixel-level correspondences between
a reference image and 2D sewing panels while transferring textures
that are free from baked-in artifacts and geometric distortions?
In this paper, we propose OmniFabric to bridge this gap. We first employ a powerful image generation model (Google, 2025) to repose the input clothing into a canonical A-pose, aligning its silhouette with the frontal render of the garment mesh simulated from the predicted sewing patterns. We further harness the temporal and spatial consistency of large-scale video generation models (Google DeepMind, 2025) to hallucinate coherent, multi-view observations of this A-pose garment. Projecting these globally consistent multi-view images onto the unwrapped sewing patterns establishes a complete, albeit coarse, texture initialization that preserves the garment’s global design. Because this initial projection inevitably bakes visual artifacts such as geometric distortions and physics-induced wrinkles into the texture map, we introduce a subsequent refinement stage operating entirely within the sewing pattern space. We design an automatic data engine that curates an extensive and highly diverse synthetic dataset of textured sewing patterns, paired with simulated texture initializations containing baked-in artifacts and distortions. With the curated dataset, we train a specialized Diffusion Transformer (DiT) that acts as a texture normalizer in the 2D sewing pattern space. This model rectifies distortions, and strips away transient illumination to produce a clean and normalized texture map. Finally, we conduct comprehensive experiments to demonstrate the superiority of OmniFabric over existing 3D texturing baselines. In summary, the key contributions of our work are:
- •
OmniFabric, a novel framework for synthesizing high-fidelity, globally coherent textures from a single in-the-wild image for 3D garment reconstruction.
- •
An automated data engine for creating realistic, diverse textured sewing patterns at scale, providing training data for learning complex garment textures.
- •
A strategy of using a canonical 3D mesh as spatial anchor to map image pixels to sewing pattern UV space, enabling globally aligned texture synthesis.
- •
A coarse-to-fine texture generation pipeline that initializes globally aligned textures, and further rectifies distortion and baked-in artifacts with a DiT-based UV space normalization network.
Remark.
Generating 3D garment textures is not a new topic. While many prior works mainly focus on directly texturing 3D garment meshes, they often do not scale well to large collections of in-the-wild images and remain weakly connected to real-world garment production pipelines. We argue that fabric texture synthesis should be grounded in the UV space of sewing patterns. This enables the model to maintain strong global coherence while remaining closely aligned with how garments are constructed in practice.
2. Related Work
2.1. 3D Garment Reconstruction
Early approaches for garment reconstruction primarily relied on implicit functions (Saito et al., 2019; Saito et al., 2020; Xiu et al., 2022) to approximate surfaces, yet they often struggle with complex topologies and loose-fitting clothing. More recent works leverage advanced representations like 3DGS (Rong et al., 2025) and image-to-3D diffusion priors (Wang et al., 2025; Luo et al., 2024) to achieve impressive visual fidelity. However, these methods typically produce rigid meshes or unstructured point clouds that are inherently unsuitable for physical cloth simulation. While methods like Garment3DGen (Sarafianos et al., 2025) bridge this gap by utilizing template deformation to produce simulation-ready assets, they rely on a slow, computationally intensive per-asset optimization process. To obtain simulation-ready assets, a significant line of work attempts to infer 2D sewing patterns directly from 3D point clouds or images. NeuralTailor (Korosteleva and Lee, 2022) focused on reconstructing pattern structures from 3D geometry, and subsequent works (Liu et al., 2023; Chen et al., 2024) predict these patterns directly from single-view images using discriminative transformer architectures. To facilitate more structured generation, GarmentCode (Korosteleva and Sorkine-Hornung, 2023) introduced a programmatic domain-specific language (DSL) for sewing patterns, enabling parametric control. Building on this representation, subsequent research has increasingly focused on advanced generative modeling. DressCode (He et al., 2024) synthesizes novel patterns from text prompts using a GPT-based framework, while AIpparel (Nakayama et al., 2025) scales into a multimodal foundation model that handles complex pattern generation natively from mixed inputs. To further improve physical realism, Dress-1-to-3 (Li et al., 2025b) incorporates differentiable physics simulators to optimize geometric alignment of the inferred patterns. Concurrently, other state-of-the-art frameworks directly predict these programmatic structures from images leveraging large generative models (Bian et al., 2025; Zhou et al., 2025; Li et al., 2025a). While these methods have established sewing patterns as a robust domain for 3D animation, the synthesis of normalized, artifact-free textures for these panels remains an open challenge.
2.2. Generative 3D Texturing
The rise of large-scale vision-language models (Radford et al., 2021; Rombach et al., 2022; Saharia et al., 2022) has revolutionized 3D texture synthesis. Early optimization-based approaches either leverage CLIP (Radford et al., 2021) for optimizing texture maps of 3D models (Hong et al., 2022; Chen et al., 2022; Michel et al., 2022; Mohammad Khalid et al., 2022), or employ Score Distillation Sampling (SDS) to optimize texture maps by distilling gradients from a frozen 2D diffusion model (Poole et al., 2022; Metzer et al., 2023; Chen et al., 2023b). While capable of generating coherent global structures, these methods tend to hallucinate generic textures without high-frequency details. To improve fidelity, recent research has shifted toward projection-based inpainting, by utilizing depth-conditioned Stable Diffusion to progressively paint 3D meshes from multiple viewpoints (Richardson et al., 2023; Chen et al., 2023a; Cao et al., 2023). These frameworks iteratively project the current texture into screen space, inpaint missing regions using 2D priors, and project the result back. More recent systems like Paint3D (Zeng et al., 2024), MVPaint (Cheng et al., 2025) and Meta 3D TextureGen (Bensadoun et al., 2024) further refine this process by synchronizing multi-view generation or employing coarse-to-fine UV refinement to minimize seams. Despite their popularity, a fundamental limitation persists across these general-purpose methods: they rely on priors trained on natural photography. Consequently, they inherently entangle illumination with surface appearance, inevitably “baking in” transient lighting effects—such as cast shadows and specular highlights—directly into the UV map. In addition, the generated textures are subject to distortion and blurring due to the complex surface topology, leading to suboptimal results.
In the specific area of garment texturing, FabricDiffusion (Zhang et al., 2024) addresses some of these pitfalls by treating texture synthesis as a material extraction task. It utilizes a diffusion model to extract tileable, distortion-free material patches from a single image without lighting artifacts. However, because it relies on local texture tiling, FabricDiffusion fails to capture global structural information, such as specific graphic placements and non-uniform texture patterns across different UV islands. OmniFabric bridges this gap by shifting from local material extraction to global texture synthesis. We operate in the canonical UV space to disentangle base-color from illumination, synthesizing a holistic texture map that respects the global texture design in the reference image. While previous methods (Zeng et al., 2024; Bensadoun et al., 2024) also deploy a 2D refinement stage on the UV map, they only perform inpainting on occluded regions and fail to rectify distortion or remove baked-in artifacts, which are the domain specific challenges for garment texturing, as shown in Figure 2.
3. Method
Figure 3 presents an overview of our approach. We aim to reconstruct 3D garments with coherent textures from real-world images. We define the task and objective in Section 3.1. Next, we introduce a two-stage pipeline for generating holistic textures directly onto sewing patterns. This includes a coarse texture map initialization (Section 3.2) and a UV space texture normalization (Section 3.3). To train the model, we develop an automated pipeline to synthesize a large scale dataset of textured sewing patterns (Section 4).
3.1. Problem Definition
Given a single reference image and an estimated 3D garment mesh parameterized by sewing patterns from an off-the-shelf sewing pattern prediction model, our goal is to generate a normalized texture map directly within the sewing pattern space. Unlike previous methods (Zhang et al., 2024) that focus on extracting local and repeatable material patches, we formulate our objective as a global texture synthesis task. We define as a representation of the normalized texture map of the garment, which is completely free of distortion, view-dependent illumination, cast shadows, and geometry-induced wrinkles. Following the paradigm of diffusion models, we formulate the synthesis of this holistic texture map as a conditional distribution mapping problem. We seek to learn a mapping function such that:
| (1) |
We note that although is not intrinsic albedo, it’s a close approximation and can be directly imported into cloth simulation engines, e.g., CLO (CLO Virtual Fashion Inc., 2024), with externally specified material assumption for the remaining properties including roughness, metallic and normal map, enabling the synthesis of simulation-ready 3D garments.
3.2. Coarse Texture Map Initialization
To guide texture synthesis in the sewing pattern space, we first aim to obtain a coarse texture initialization that fully leverages the visual information from the reference image. This requires establishing pixel-wise correspondences between the reference image and the sewing pattern. While this mapping is challenging, we recast the task via a VLM-driven and pose aware texture alignment step. We further leverage the multi-view generative priors of large video models to provide a complete and globally coherent texture initialization.
3.2.1. Pose aware texture alignment
Using Nano Banana Pro (Google, 2025), we repose the reference image into to perfectly align with the frontal silhouette of the rest-pose garment mesh . This alignment enables direct pixel projection of onto and its corresponding sewing patterns , producing the textured frontal render . We observe that Nano Banana Pro successfully manipulates the input garment into a desired pose while preserving the original texture with high accuracy and fidelity.
3.2.2. Consistent multi-view synthesis.
After the initial pose alignment, a significant portion of the UV space occluded from the frontal view remains untextured. To hallucinate these missing regions while encouraging strong spatial continuity, we leverage a large-scale video generation model for multi-view generation. By conditioning the model on the textured frontal render , we synthesize a 360-degree spinning video of the garment, naturally exploiting the model’s inherent temporal priors for structural consistency. Following recent projection-based texturing paradigms, we sample multi-view , four orthogonal views (front, back, left, and right) from the generated video. These observations are then projected onto and mapped to their corresponding UV coordinates on to form the coarse texture map .
3.3. UV Space Texture Refinement
While preserves the coherent global texture of the garment, it contains baked-in artifacts like shadows, physical wrinkles and unpainted areas. Furthermore, spatial distortions in leave it far short of production-ready fabric quality, as shown in Figure 2. To address these issues, we introduce a texture normalization model designed to rectify into a normalized texture map .
3.3.1. Training objective of multi-condition diffusion
We formulate texture normalization as a holistic UV synthesis problem by developing a conditional distribution mapping network. Unlike FabricDiffusion (Zhang et al., 2024), which relies on a single condition, i.e., a local textile patch, for the normalization task, our framework incorporates a multimodal conditioning set to guide global texture normalization. We extend Eq. 1 and minimize the following:
| (2) |
where represents the noisy latent of the ground-truth texture at timestep . The multi-modal conditioning set includes the coarsely-textured map , textured frontal rendering and a 3D position map , created by projecting the vertices’ 3D positions of onto the 2D layout of . helps the model learn spatial connectivity between separate UV panels, ensuring seamless textures across fragmented UV islands.
3.3.2. Model architecture and training
To process multimodal conditioning set , we adopt a Diffusion Transformer (DiT) architecture following previous works (Tan et al., 2025a; Tan et al., 2025b), and leverage a unified token-processing strategy that treats all conditions as a unified sequence of tokens. We utilize a joint self-attention mechanism where the condition tokens (from , , and ) and the noisy latent tokens from interact within the same transformer blocks. Because , , and the target all share the same 2D panel layout, we apply a dynamic positional encoding strategy: image patches at the same spatial coordinates across these maps are assigned identical positional encodings, while only retains its own coordinate system through distinct encodings. This minimal yet universal design allows the model to attend to the structural cues of while simultaneously hallucinating missing textures in . We use a pretrained weight of DiT, and fine-tune the model on our synthetic textured sewing pattern data with LoRA (Hu et al., 2022).
4. Synthetic Training Data Creation
To train our texture normalization model, we develop a data engine (Figure 4) to curate a large-scale dataset of textured sewing patterns. The core objective is to synthesize training pairs that exhibit the complex appearance of real-world garments while providing ground-truth normalized texture maps for supervised learning.
4.1. Textured Sewing Pattern Creation
A key bottleneck in garment dataset construction is the lack of complex and high-resolution textures; most available texture sources provide only simple, repetitive patterns. To overcome this, we introduce a pipeline to produce diverse, high-fidelity texture images by using fashion images from real-world datasets, e.g., DeepFashion (Liu et al., 2016), as references. For each sample, we apply a Large Language Model (LLM) (Bai et al., 2023) to generate a detailed description encompassing texture layout, semantic elements, and color palettes. These descriptions, combined with the reference image, are fed into a vision-language model (Google, 2025) to synthesize textures with diverse designs. We then overlay the generated texture image onto sewing patterns sampled from GarmentCodeData. Note that the goal here is not to reconstruct the exact texture from the example image, but to generate plausible and realistic textures that provide sufficient diversity for training the texture synthesis model.
4.2. Training Pairs Construction
For each textured sewing pattern, we generate a training tuple consisting of a coarse texture map, a frontal rendering, a positional map, and the ground-truth texture. To explicitly train the network to rectify geometric distortions, we apply a random Thin Plate Spline (TPS) (Bookstein, 1989) distortion to the initial texture image. The original, undistorted texture forms the ground-truth map , while the TPS-distorted texture is overlaid onto the sewing pattern for physical simulation. Specifically, we perturb each control point of a regular grid fitted to the texture’s aspect ratio, giving with
| (3) |
where by default controls the overall distortion magnitude and is the resolution of a texture image. A TPS is then fit between the regular grid and its perturbed counterpart and applied densely to warp the full-resolution texture. We stitch and drape the sewing pattern onto an A-pose SMPL body to yield a 3D garment mesh , from which we capture the frontal rendering . The positional map is then derived by projecting the coordinates of the mesh vertices onto the 2D panel layout . Finally, we synthesize the coarse texture map by projecting and a randomly sampled view onto a neutral-posed mesh and unwrapping them back into the 2D pattern space. This explicitly bakes physics-induced artifacts and deformations into the initialization, compelling the model to learn undistortion and normalized base-color recovery.
5. Experiments
In this section, we present a comprehensive evaluation of OmniFabric through quantitative and qualitative analyses. We first detail the experimental setup, including curated synthetic dataset and metrics used for evaluation. We then demonstrate the effectiveness of our framework by comparing against SOTA methods for 3D texturing. Finally, we conduct ablation studies to validate our core architectural designs. Additional implementation details, extended qualitative results, our method’s adaptability to other sewing pattern prediction methods besides ChatGarment (Bian et al., 2025), examples of our curated garment data, and discussion on potential future directions are provided in the supplementary material.
5.1. Setup
5.1.1. Datasets
To train and evaluate our model, we construct a large-scale synthetic dataset of textured sewing patterns, as described in Section 4. We sample a total of 3K unique and diverse garment samples from GarmentCodeData (Korosteleva et al., 2024). We then use the pipeline in Figure 4 to generate texture images, and print them onto sampled sewing patterns, finally constructing a dataset of 30K textured sewing patterns with diverse appearances and structural designs (10 texture variations per garment style on average). We simulate garments using NVIDIA Warp (Macklin, 2022) with a per-sample random seed, so repeated simulations of the same sewing pattern produce distinct meshes and, consequently, different baked-in wrinkles in . For lighting, we sample random subsets of white point lights to introduce varied illumination and baked-in shading across the dataset. We train our model with this synthetic dataset, and keep a held-out set for quantitative comparison with other baseline models with a test/train ratio of 0.1. Besides synthetic data, we evaluate all methods on in-the-wild images, which is the major focus of this work. These in-the-wild images are obtained from DeepFashion (Liu et al., 2016) dataset and recent clothing-design images collected from the web. Due to the lack of ground truth in-the-wild images, we evaluate mainly through qualitative comparisons.
5.1.2. Metrics
We evaluate fidelity and structural coherence of the generated textures with a suite of standard metrics. We use LPIPS (Zhang et al., 2018) and DISTS (Ding et al., 2020) to measure visual similarity and structural consistency. We also compare with SSIM (Wang et al., 2004) and MS-SSIM (Wang et al., 2003), metrics employed to assess pixel-level structural accuracy. CLIP-score (CLIP-s) (Gal et al., 2022) measures the semantic alignment between the final textured garment and the reference images.
5.1.3. Baseline methods
We compare against several SOTA methods in 3D texturing and garment-specific texture transfer. Among them, Hunyuan3D-2.0 (Zhao et al., 2025) and Paint3D (Zeng et al., 2024) use multi-view generation to provide additional prior beyond the given single-view observation. FabricDiffusion (Zhang et al., 2024) is capable of creating fabric textures with no baked-in artifacts by normalizing local textile patterns with a diffusion model.
5.1.4. Implementation details
Our texture normalization model is based on the Diffusion Transformer (DiT) architecture, and we follow the multi-conditioning strategy of OminiControl (Tan et al., 2025a) by treating tokens from all input conditions as a unified token sequence. The model is fine-tuned using LoRA on a single NVIDIA A6000 GPU with a batch size of 8, keeping the VAE and primary DiT backbone fixed to preserve generative stability. We resize and pad sewing pattern maps to a resolution of 10241024 in a batch for training efficiency. During inference, we reverse the padding and resizing process, and deploy an image super-resolution model (Wang et al., 2021) to upscale the output. Since we are fine-tuning the model with LoRA, a higher training resolution can be easily achieved without GPU memory issues, making our framework scalable for training data of higher resolution. We fine-tune the model until convergence on our synthetic dataset. We utilize Veo 3 (Google DeepMind, 2025) as the multi-view video generation prior and Nano Banana Pro (Google, 2025) for the initial texture transfer to the frontal rendering . We used ChatGarment (Bian et al., 2025) as the model for sewing pattern prediction throughout all experiments in the main paper. The average runtime for a single inference is roughly 5 minutes, depending on the current load on the Gemini model servers. The failure rate of Gemini models, e.g., generating unrelated textures on or multi-views, is lower than 2%, estimated from randomly sampled generations. No additional filtering is necessary thanks to its stable performance. Additional details including prompts used for generation are provided in supplementary material.
5.2. Qualitative Comparisons
5.2.1. Texture synthesis by OmniFabric from in-the-wild images
We first show our results in Figure 5. Given a reference image, OmniFabric first generates a coarse texture map, serving as a roughly textured sewing pattern. Then, the texture normalization model removes the distortions and baked-in artifacts. From the front views of simulated 3D garments provided in Figure 5, we show that OmniFabric is capable of creating seamless textures across the 3D surface of garments. Note that while OmniFabric excels in textured garment synthesis from real-world images, we fine-tune our model only with synthetic data, which poses great scalability for model improvement.
5.2.2. Simulation results
We present simulations of synthesized 3D garments in Figure 9 via CLO. The sewing patterns generated by OmniFabric provide the normalized RGB base-color as albedo. Other parameters not predicted by our model—roughness, metallic, reflection intensity, and auto-generated normal map—are left at CLO’s default Fabric_Matte preset values. The Environment/HDRI maps for relighting examples are also chosen from CLO’s lighting presets. The simulated assets exhibit a convincing and photorealistic dynamic appearance, maintaining structural integrity and texture coherence even under extreme dynamics and varying illumination.
5.2.3. Comparison with state-of-the-art methods
We compare with other methods in Figure 6 using in-the-wild images as well. For input with complex textures and elements, both Hunyuan3D and Paint3D fail to preserve the high-frequency details. Even though FabricDiffusion can create normalized textile pattern and texture the given mesh by tiling local textile patches, it is limited to local textile patches and fails to generate coherent global appearance. Conversely, OmniFabric creates normalized textures while preserving coherent details.
5.3. Quantitative Comparisons
5.3.1. Synthetic dataset
We show quantitative results where all baselines are evaluated with our curated synthetic dataset for fair comparison. For each test case, we provide every baseline with the identical frontal rendering and rest-posed mesh . Each method then performs its respective 3D texturing task on . To compute image-based metrics, the resulting 3D textures are projected back into the sewing pattern UV space to generate a comparable texture map for each baseline. Table 1 shows OmniFabric excels in all metrics, particularly in preserving global structural coherence and eliminating projection artifacts.
5.3.2. Real-world data
| LPIPS | SSIM | MS-SSIM | DISTS | CLIP-s | |
|---|---|---|---|---|---|
| FabricDiffusion (Zhang et al., 2024) | 0.273 | 0.645 | 0.707 | 0.259 | 0.906 |
| Paint3D (Zeng et al., 2024) | 0.311 | 0.657 | 0.704 | 0.266 | 0.890 |
| Hunyuan3D-2.0 (Zhao et al., 2025) | 0.223 | 0.712 | 0.717 | 0.220 | 0.924 |
| OmniFabric (ours) | 0.092 | 0.868 | 0.905 | 0.121 | 0.963 |
| Overall rank | Fidelity | Back-view | Global Coherence | |
|---|---|---|---|---|
| FabricDiffusion (Zhang et al., 2024) | 3.40 | 1.75 | 2.69 | 3.15 |
| Paint3D (Zeng et al., 2024) | 3.41 | 1.86 | 2.81 | 3.03 |
| Hunyuan3D-2.0 (Zhao et al., 2025) | 2.10 | 3.43 | 3.41 | 3.50 |
| OmniFabric (ours) | 1.09 | 4.44 | 4.35 | 4.41 |
Besides comparing with other methods in Figure 6, we conduct a user study to evaluate the performances of the methods due to the absence of real-world dataset of textured sewing patterns. As detailed in Table 2, the study involves 13 participants evaluating results from 10 real-world image inputs based on four criteria: overall quality ranking (1 to 4, lower is better), fidelity to the input appearance, back-view plausibility, and global texture coherence (1 to 5, higher is better). OmniFabric achieves the best results across all criteria, demonstrating that it is preferred on real images and more faithfully preserves coherent garment textures.
5.4. Reproducibility via Open-Source Models
| LPIPS | SSIM | DISTS | CLIP-s | |
|---|---|---|---|---|
| OmniFabric w/ default settings | 0.092 | 0.868 | 0.121 | 0.963 |
| OmniFabric w/ open-source models | 0.114 | 0.847 | 0.149 | 0.955 |
| Replacement Setting | LPIPS | SSIM | DISTS | CLIP-s |
|---|---|---|---|---|
| (a) GT + GT | 0.105 | 0.849 | 0.159 | 0.944 |
| (b) Pred. + GT | 0.121 | 0.837 | 0.176 | 0.935 |
| (c) Pred. + Pred. | 0.144 | 0.804 | 0.166 | 0.938 |
| (d) Pred. + Pred. + SP Norm. | 0.092 | 0.868 | 0.121 | 0.963 |
| Setting | LPIPS | SSIM | DISTS | CLIP-s |
|---|---|---|---|---|
| OmniFabric (ours) | 0.092 | 0.868 | 0.121 | 0.963 |
| (a) w/o coarse texture map | 0.407 | 0.601 | 0.306 | 0.887 |
| (b) w/o frontal view rendering | 0.097 | 0.850 | 0.129 | 0.962 |
| (c) w/o position map | 0.098 | 0.848 | 0.130 | 0.961 |
While using Gemini models for pose transfer and multi-view generation (Section 3.2) as default implementation of OmniFabric, we further construct a fully open-source alternative to facilitate reproducibility. To replace the closed-source models, we use FLUX.2 (Labs, 2025) as image generation backbone and incorporate Ministral 3 (Liu et al., 2026) for prompt up-sampling. We feed textured frontal view and silhouettes of four orthogonal views as conditions to generate the multi-view with this alternative. Table 3 shows that open-source alternative achieves similar performance with the same experiment settings in Section 5.3.
5.5. Error-Accumulation Analysis
We analyze error accumulation across modules in our model architecture in this section by progressively replacing intermediate outputs, specifically the reposed reference image and generated multi-view , with ground-truth counterparts. As shown in Table 4, settings (a) to (c) accumulate errors across stages including texture projection, pose alignment and multi-view generation, and (d) shows that our normalization model effectively fixes these aggregated artifacts on sewing patterns.
5.6. Ablation Studies and Analyses
5.6.1. Importance of the canonical 3D and coarse texture map
As shown in Table 5-(a), without the coarse texture map, which serves as the distillation of generative priors, the performance drops significantly. This shows that the projected textures do provide essential guidance for the texture generation. Figure 8 shows that the model is randomly generating textures that “look similar” instead of pixel-aligned content without coarse texture map .
5.6.2. Effect of the frontal view rendering condition
We observe that removing during training leads to inferior results (Table 5-(b)), suggesting that the frontal-view rendering provides important cues for the model. In particular, offers a holistic visual reference of the garment appearance, helping the model better understand global layout and consistency of the texture across sewing pattern pieces.
5.6.3. Effect of position map
5.6.4. Importance of geometric distortion in dataset.
The right side of Figure 8 shows the necessity of creating geometric distortion in the synthetic training data. Through training to remove geometric distortion, the model can not only generate normalized structured patterns on the sleeves, e.g., the checker pattern, but also produce clean appearances for other elements, e.g., texts and logos.
5.7. Limitations
As shown in Figure 7, correctness of synthesized textures can be affected by inaccurately generated multi-views (1st row). OmniFabric is also unable to refine the mismatched geometry predicted by ChatGarment (Bian et al., 2025) (2nd row); however, it does not change the role of our refinement model: given a sewing pattern, it normalizes the textures by reducing distortions, wrinkles, shadows, and projection artifacts, regardless of specific garment geometry. We put more discussion of limitations in the supplementary material.
6. Conclusion and Discussion
We introduce OmniFabric, a novel framework for synthesizing 3D garments with normalized and globally coherent textures. By leveraging large-scale generative priors and training a texture normalization model, we effectively disentangle intrinsic albedo from view-dependent artifacts including baked-in shadows and wrinkles. Our method bridges the gap between single-view observations and simulation-ready 3D assets, maintaining structural integrity across complex sewing patterns. For this study, we focus on the synthesis of sewing patterns with normalized textures within the GarmentCode framework. An immediate and promising extension would be to consider the joint prediction of complete PBR material maps, including roughness, metallic, and normal maps, alongside the albedo. We believe this direction will further enhance the photorealism of the reconstructed garments under diverse environmental lighting, providing even greater utility for immersive digital content creation.
References
- Qwen technical report. arXiv preprint arXiv:2309.16609. Cited by: §4.1.
- Meta 3d texturegen: fast and consistent texture generation for 3d objects. arXiv preprint arXiv:2407.02430. Cited by: §1, §2.2, §2.2.
- Chatgarment: garment estimation, generation and editing via large language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 2924–2934. Cited by: §1, §2.1, §5.1.4, §5.7, §5.
- Principal warps: thin-plate splines and the decomposition of deformations. IEEE Transactions on pattern analysis and machine intelligence 11 (6), pp. 567–585. Cited by: §4.2.
- Texfusion: synthesizing 3d textures with text-guided image diffusion models. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 4169–4181. Cited by: §2.2.
- Panelformer: sewing pattern reconstruction from 2d garment images. In Proceedings of the IEEE/CVF winter conference on applications of computer vision, pp. 454–463. Cited by: §2.1.
- Text2tex: text-driven texture synthesis via diffusion models. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 18558–18568. Cited by: §1, §2.2.
- Fantasia3d: disentangling geometry and appearance for high-quality text-to-3d content creation. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 22246–22256. Cited by: §2.2.
- Tango: text-driven photorealistic and robust 3d stylization via lighting decomposition. Advances in neural information processing systems 35, pp. 30923–30936. Cited by: §2.2.
- Mvpaint: synchronized multi-view diffusion for painting anything 3d. In Proceedings of the Computer Vision and Pattern Recognition Conference, pp. 585–594. Cited by: §2.2.
- CLO. Note: https://www.clo3d.comVersion 7.2 Cited by: §3.1.
- Image quality assessment: unifying structure and texture similarity. IEEE transactions on pattern analysis and machine intelligence 44 (5), pp. 2567–2581. Cited by: §5.1.2.
- An image is worth one word: personalizing text-to-image generation using textual inversion. arXiv preprint arXiv:2208.01618. Cited by: §5.1.2.
- Veo: google’s most capable video generation model. Note: https://deepmind.google/models/veo/Accessed: September 24, 2026 Cited by: §1, §5.1.4.
- Gemini 3 pro image. Note: https://gemini.google.com/ Cited by: §1, §3.2.1, §4.1, §5.1.4.
- Real-time deep dynamic characters. ACM Transactions on Graphics (ToG) 40 (4), pp. 1–16. Cited by: §1.
- Dresscode: autoregressively sewing and generating garments from text guidance. ACM Transactions on Graphics (TOG) 43 (4), pp. 1–13. Cited by: §1, §2.1.
- Avatarclip: zero-shot text-driven generation and animation of 3d avatars. arXiv preprint arXiv:2205.08535. Cited by: §2.2.
- Lora: low-rank adaptation of large language models.. Iclr 1 (2), pp. 3. Cited by: §3.3.2.
- Physical simulation for animation and visual effects: parallelization and characterization for chip multiprocessors. ACM SIGARCH Computer Architecture News 35 (2), pp. 220–231. Cited by: §1.
- Effects of 3d virtual “try-on” on online sales and customers’ purchasing experiences. Ieee Access 8, pp. 189479–189489. Cited by: §1.
- Consumer experience in using 3d virtual garment simulation technology. Journal of the Textile Institute 104 (8), pp. 819–829. Cited by: §1.
- GarmentCodeData: a dataset of 3d made-to-measure garments with sewing patterns. In European Conference on Computer Vision, pp. 110–127. Cited by: §5.1.1.
- Neuraltailor: reconstructing sewing pattern structures from 3d point clouds of garments. ACM Transactions on Graphics (TOG) 41 (4), pp. 1–16. Cited by: §2.1.
- Garmentcode: programming parametric sewing patterns. ACM Transactions on Graphics (TOG) 42 (6), pp. 1–15. Cited by: §2.1.
- FLUX.2: Frontier Visual Intelligence. Note: https://bfl.ai/blog/flux-2 Cited by: §5.4.
- GarmentDiffusion: 3d garment sewing pattern generation with multimodal diffusion transformers. arXiv preprint arXiv:2504.21476. Cited by: §2.1.
- Dress-1-to-3: single image to simulation-ready 3d outfit with diffusion prior and differentiable physics. ACM Transactions on Graphics (TOG) 44 (4), pp. 1–16. Cited by: §1, §2.1.
- Ministral 3. arXiv preprint arXiv:2601.08584. Cited by: §5.4.
- Towards garment sewing pattern reconstruction from a single image. ACM Transactions on Graphics (TOG) 42 (6), pp. 1–15. Cited by: §2.1.
- DeepFashion: powering robust clothes recognition and retrieval with rich annotations. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: §4.1, §5.1.1.
- Garverselod: high-fidelity 3d garment reconstruction from a single in-the-wild image using a dataset with levels of details. ACM Transactions on Graphics (TOG) 43 (6), pp. 1–12. Cited by: §2.1.
- Warp: A High-performance Python Framework for GPU Simulation and Graphics. Note: NVIDIA GPU Technology Conference (GTC) External Links: Link Cited by: §5.1.1.
- Latent-nerf for shape-guided generation of 3d shapes and textures. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 12663–12673. Cited by: §2.2.
- Text2mesh: text-driven neural stylization for meshes. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 13492–13502. Cited by: §2.2.
- Clip-mesh: generating textured meshes from text using pretrained image-text models. In SIGGRAPH Asia 2022 conference papers, pp. 1–8. Cited by: §2.2.
- Aipparel: a multimodal foundation model for digital garments. In Proceedings of the Computer Vision and Pattern Recognition Conference, pp. 8138–8149. Cited by: §C.2, §1, §2.1.
- Dreamfusion: text-to-3d using 2d diffusion. arXiv preprint arXiv:2209.14988. Cited by: §2.2.
- Learning transferable visual models from natural language supervision. In International conference on machine learning, pp. 8748–8763. Cited by: §2.2.
- Texture: text-guided texturing of 3d shapes. In ACM SIGGRAPH 2023 conference proceedings, pp. 1–11. Cited by: §1, §2.2.
- High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 10684–10695. Cited by: §2.2.
- Gaussian garments: reconstructing simulation-ready clothing with photorealistic appearance from multi-view video. In 2025 International Conference on 3D Vision (3DV), pp. 1054–1063. Cited by: §1, §2.1.
- Photorealistic text-to-image diffusion models with deep language understanding. Advances in neural information processing systems 35, pp. 36479–36494. Cited by: §2.2.
- Pifu: pixel-aligned implicit function for high-resolution clothed human digitization. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 2304–2314. Cited by: §2.1.
- Pifuhd: multi-level pixel-aligned implicit function for high-resolution 3d human digitization. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 84–93. Cited by: §2.1.
- Garment3dgen: 3d garment stylization and texture generation. In 2025 International Conference on 3D Vision (3DV), pp. 1382–1393. Cited by: §1, §2.1.
- Ominicontrol: minimal and universal control for diffusion transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 14940–14950. Cited by: §3.3.2, §5.1.4.
- Ominicontrol2: efficient conditioning for diffusion transformers. arXiv preprint arXiv:2503.08280. Cited by: §3.3.2.
- Real-esrgan: training real-world blind super-resolution with pure synthetic data. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 1905–1914. Cited by: §5.1.4.
- GarmentCrafter: progressive novel view synthesis for single-view 3d garment reconstruction and editing. arXiv preprint arXiv:2503.08678. Cited by: §1, §2.1.
- Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing 13 (4), pp. 600–612. Cited by: §5.1.2.
- Multiscale structural similarity for image quality assessment. In The Thirty-Seventh Asilomar Conference on signals, systems & computers, 2003, Vol. 2, pp. 1398–1402. Cited by: §5.1.2.
- Icon: implicit clothed humans obtained from normals. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 13286–13296. Cited by: §2.1.
- Paint3d: paint anything 3d with lighting-less texture diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 4252–4262. Cited by: §1, §2.2, §2.2, §5.1.3, Table 1, Table 2.
- FabricDiffusion: high-fidelity texture transfer for 3d garments generation from in-the-wild images. In SIGGRAPH Asia 2024 Conference Papers, pp. 1–12. Cited by: §1, §2.2, §3.1, §3.3.1, §5.1.3, Table 1, Table 2.
- The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 586–595. Cited by: §5.1.2.
- Hunyuan3d 2.0: scaling diffusion models for high resolution textured 3d assets generation. arXiv preprint arXiv:2501.12202. Cited by: §1, §5.1.3, Table 1, Table 2.
- Design2GarmentCode: turning design concepts to tangible garments through program synthesis. In Proceedings of the Computer Vision and Pattern Recognition Conference, pp. 23712–23722. Cited by: §2.1.
Supplementary Material
Supplementary Material A A Details of Dataset Construction
A primary contribution of this work is the automated data engine designed to generate high-fidelity textured sewing patterns. While high-detail manual texturing is possible, the process remains prohibitively time-consuming for large-scale applications. In contrast, our pipeline efficiently produces diverse garment styles and textures with a vast range of complexity, facilitating the construction of digital garment datasets at scale. This automated engine is a core contribution that addresses the scarcity of complex, high-resolution textures in existing garment research. We have included several representative examples in this supplement and will release the full dataset as well as the curation pipeline.
A.1. Reordering Sewing Pattern for Texturing
To align our generation process with real-world manufacturing, we introduce a reordering strategy for sewing pattern layouts. In professional garment construction, textures are often seamless across specific functional groups—such as a large frontal fabric panel—but exhibit discontinuities at structural seams, such as the transition from a sleeve to the chest. These discontinuities arise because 3D-stitched edges are often not geometrically matched in 2D space (e.g., a curved sleeve head joining a straight armhole), necessitating different fabric cuts. Following these observations, our pipeline first reorders the 2D sewing pattern pieces by merging related panels, such as the front and back pieces of a single sleeve, into unified spatial groups. By overlaying textures onto this reordered 2D layout, we ensure seamless pattern continuity within each functional group while maintaining the realistic texture breaks required for production-ready 3D assets.
Supplementary Material B B Additional Implementation Details
In this section, we provide further technical specifics regarding our texturing pipeline. As detailed in the main paper, we utilize a single-view garment observation to generate a 360-degree rotation video. From this sequence, we extract four keyframes corresponding to the orthogonal front, back, left, and right camera views for the multi-view texture projection process. Given that the video depicts a garment in a canonical A-pose, these four views are sufficient to capture the vast majority of surface details. The rest-posed mesh, , is reconstructed by stitching and draping the predicted sewing patterns onto a A-posed SMPL human body model.
B.1. Non-overlapping Texture Projection
A key distinction between our projection method and state-of-the-art baselines, such as Hunyuan3D-2.0, lies in the source and application of multi-view data. While traditional methods project independent images generated by multi-view diffusion models, we project selected frames from a structurally consistent video prior using a specialized non-overlapping strategy. We observe that while multi-view diffusion models improve cross-view alignment, they often sacrifice high-frequency details to maintain that consistency. To preserve these details, our non-overlapping strategy begins by projecting textures from the front and back views—the two perspectives with minimal observational overlap—directly onto the mesh. Subsequently, textures from the side views are projected exclusively onto previously untextured regions, effectively performing a projection-based inpainting. Finally, our data-driven texture normalization model rectifies any minor artifacts or discontinuities at the seams of these projected textures. This approach allows OmniFabric to retain intricate, high-frequency surface details while ensuring global multi-view consistency.
B.2. Prompt Configurations
In this section we specify the design of prompts for the Gemini models in our framework. We use the prompt below for pose aware texture alignment, which generates a textured frontal view of rest-pose garment mesh :
The first apparel is my target apparel. Please generate textures on this apparel, so the textures are identical to the apparel in the reference image. Please don’t change the style, size and pose of the target apparel.
For the multi-view generation with video model, we use the following prompt:
Create a continuous 360 degree rotation video of this apparel, making it rotate horizontally 360 degree, like a microwave.
B.3. Training Details
We fine-tune all three baselines from pretrained weights, adapting the components described below. For Paint3D, we fine-tune its position encoder, as specified in its paper. For Hunyuan3D, we fine-tune its multi-view image generator. And for FabricDiffusion, we train its texture generator using paired textile data. All the models are fine-tuned upon pre-trained model weights.
Supplementary Material C C Additional Results and Analyses
C.1. Robustness to Structural Variations in Sewing Patterns
We present further qualitative evaluations in Figure S3, which demonstrate the robustness of our framework across a variety of garment styles. The results illustrate that our texture normalization model effectively rectifies the coarse texture map by eliminating projection-induced distortions and removing physics-based wrinkles inherent in the initial capture. While the geometry simulated from the predicted sewing patterns may occasionally deviate from the exact silhouette in the input image, we emphasize that such discrepancies arise from the limitations of the underlying sewing pattern prediction method rather than the texturing process. To mitigate this, our pipeline employs a texture transferring model that aligns the input image textures with the rendered silhouette of the predicted garment. This strategy successfully bridges the gap between the predicted geometry and the original observation, ensuring that the synthesized textures remain globally coherent and structurally aligned.
C.2. Compatibility with Other Sewing Pattern Prediction Method
While the main paper utilizes ChatGarment for sewing pattern prediction, our framework is designed to be agnostic to the specific reconstruction method employed. To demonstrate this flexibility, we evaluated our pipeline using sewing patterns predicted by Alpparel (Nakayama et al., 2025), a recent state-of-the-art approach based on Large Vision-Language Models. As illustrated in Figure S2, OmniFabric consistently generates textured garments with aligned geometry and high-frequency details, confirming that our synthesis capability is not restricted to a single pattern reconstruction architecture.It should be noted, however, that the simulated garment geometry may occasionally exhibit minor mismatches relative to the original input image due to imperfect pattern prediction. As discussed in Section C.1, our methodology specifically focuses on maintaining global texture consistency even when the underlying 3D mesh, which is simulated from predicted sewing pattern, is not perfectly aligned with the reference observation. This robustness ensures that OmniFabric can be seamlessly integrated with future advancements in sewing pattern prediction.
Supplementary Material D D Limitations and Future Work
While OmniFabric represents a significant advancement in garment texturing, several limitations remain that offer promising avenues for future research. First, although our model effectively removes transient lighting artifacts to produce a clean appearance, it does not yet explicitly disentangle albedo from shading in a strictly principled, physics-based manner, a distinction that should be noted to avoid overstating our current appearance decomposition capabilities. Furthermore, because the framework relies on generative priors, unseen or occluded regions are occasionally hallucinated; however, the system is not strictly limited to single-view observations and could incorporate multiple frames or video sequences in practice to reduce these hallucinations and improve reconstruction fidelity. Finally, certain materials with strong view-dependent effects, such as highly reflective leather or metallic fabrics, remain challenging to reconstruct faithfully, suggesting that future iterations should incorporate the joint prediction of complete PBR material maps—including roughness and normal maps—to achieve true photorealism across diverse environmental lighting conditions.