Delivering high-quality 3D models in a web browser is one of the most technically demanding aspects of web product development. A GLB file that looks excellent in Blender or Unity may weigh 80MB, load in 12 seconds on a mobile connection and consume enough memory to crash Safari on an older iPhone. This article covers the complete pipeline for optimising GLB models for web delivery — from geometry reduction through texture compression to runtime loading strategies.
Why GLB Optimisation Matters More Than You Think
The web is a hostile environment for 3D content. Unlike native applications where you control the target hardware, web 3D must run on everything from a 2020 MacBook Pro to an Android phone from 2019 with 2GB of RAM. The browser imposes additional constraints: WebGL has a maximum texture size limit (typically 4096×4096 on mobile), JavaScript memory limits can cause tab crashes above 500MB–1GB, and the HTTP request overhead for large binary files causes significant latency even on fast connections.
The goal of GLB optimisation is not to make models look worse — it is to make them look just as good at a fraction of the original file size and memory footprint. Done correctly, an optimised GLB can be 10–20x smaller than the original Blender export with no perceptible quality difference in the final rendering.
Step 1: Geometry Reduction with Meshopt
Meshopt is the geometry compression codec built into the glTF specification and supported by Three.js, Babylon.js and most modern WebGL renderers. It works by encoding vertex data with delta coding and variable-length encoding, and reordering index buffers for better GPU cache coherency. In practice, Meshopt reduces the geometry portion of a GLB file by 40–70% with zero quality loss — the mesh is decoded to exactly the same vertices on load.
Apply Meshopt using the gltfpack command-line tool: gltfpack -i input.glb -o output.glb -cc. The -cc flag enables all compression including Meshopt. For models with dense geometry you can also use -si 0.8 to simplify geometry by 20% as part of the same pass, which is often imperceptible for background or secondary models.
Step 2: Texture Compression with KTX2 and Basis
Textures are typically 60–80% of a GLB file size and, more critically, of GPU memory consumption. A 4096×4096 RGBA PNG texture decompresses to 64MB of GPU memory. KTX2 with Basis Universal encoding solves this by providing GPU-native compressed textures that remain compressed in GPU memory. A 4096×4096 KTX2/ETC1S texture uses approximately 8MB of GPU memory — an 8x reduction.
KTX2 has two encoding modes: ETC1S (better compression ratio, slightly lower quality, ideal for albedo/diffuse textures) and UASTC (better quality, larger file size, ideal for normal maps and detail textures). Use ETC1S for base colour, roughness and metalness maps. Use UASTC for normal maps where quality degradation would cause visible artefacts in lighting calculations.
Generate KTX2 textures using the KTX-Software toktx tool or the gltf-transform CLI. In Three.js, KTX2 textures require the KTX2Loader and a GPU decoder — load the decoder WASM file from the Three.js examples directory and pass it to KTX2Loader before loading any GLB that contains KTX2 textures.
Step 3: Draco Mesh Compression
Draco is a geometry compression library that reduces mesh file size by encoding triangle connectivity and vertex attributes with lossy or lossless compression. It complements Meshopt rather than replacing it — Draco provides better compression ratios for simple geometry with regular topology, while Meshopt is faster to decode and better for complex organic shapes.
Use Draco for architectural models, furniture and product visualisation where geometry is relatively clean. Avoid Draco for character models and organic shapes where the lossy quantisation can introduce visible artefacts in curved surfaces. In Three.js, Draco requires the DRACOLoader with a decoder WASM — load it once and reuse across all GLB files.
Step 4: LOD (Level of Detail) Strategy
Even with compression, loading all geometry at full resolution for objects that appear small on screen is wasteful. LOD systems swap high-resolution geometry for lower-resolution variants based on screen space coverage. In Three.js, the LOD object handles this automatically based on camera distance.
Generate LOD meshes using the Decimate modifier in Blender or the simplify command in gltfpack. A typical LOD chain has three levels: LOD0 at full resolution, LOD1 at 50% triangle count and LOD2 at 15% triangle count. For large scenes with many objects, LOD alone can reduce average draw call geometry by 60–70%.
Step 5: Progressive Loading and Streaming
Even a well-optimised 5MB GLB file should not block the initial page render. We use progressive loading strategies to show the page immediately while 3D content loads in the background. The pattern is: render a poster image (JPEG/WebP screenshot of the 3D content) immediately, load a low-poly placeholder GLB (< 100KB) first, then lazy-load the full-quality model when the user interacts with the 3D viewer or when the page is idle.
For scenes with many objects, use GLTF's built-in extension for sparse accessors and load critical objects first. The Three.js GLTFLoader fires progress events during load — use these to update a loading bar and swap from poster to 3D as soon as the first mesh is available.
Practical Results: Before and After
A typical unoptimised product visualisation GLB from a design team might weigh 45MB with PNG textures. After applying Meshopt, KTX2/ETC1S texture compression and a one-level LOD for background objects, the same model typically comes in at 3–6MB — a 7–15x reduction. GPU memory consumption drops from 280MB to 35–50MB, making it viable on mobile devices that were previously unable to load the model at all.
Key Takeaways
- Apply Meshopt compression to all geometry — it is lossless and reduces file size by 40–70% with no quality impact.
- Use KTX2/ETC1S for albedo textures and KTX2/UASTC for normal maps — both stay compressed in GPU memory.
- Implement LOD for scenes with more than 20 distinct objects — it significantly reduces per-frame geometry cost.
- Never block page render on 3D content — use poster images and progressive loading to keep Time to Interactive low.
- Test on mobile hardware during development — GPU memory limits on mobile are the most common source of production failures.


