# Real-ESRGAN models for UtilPick "AI 화질 개선" (image-upscale)

These ONNX files are conversions of the official Real-ESRGAN release weights.
They run in the visitor's browser (onnxruntime-web); images are never uploaded.

Real-ESRGAN — Copyright (c) 2021, Xintao Wang. BSD 3-Clause License (see `LICENSE` in this folder).
Project: https://github.com/xinntao/Real-ESRGAN · Paper: Wang et al., "Real-ESRGAN: Training Real-World Blind
Super-Resolution with Pure Synthetic Data", ICCVW 2021.

## Files

| File | Source weights | Architecture | Size (bytes) | SHA-256 |
| --- | --- | --- | --- | --- |
| `realesr-general-x4v3-fp16.onnx` | [realesr-general-x4v3.pth](https://github.com/xinntao/Real-ESRGAN/releases/download/v0.2.5.0/realesr-general-x4v3.pth) (release v0.2.5.0, 4,885,111 bytes, sha256 `8dc7edb9ac80ccdc30c3a5dca6616509367f05fbc184ad95b731f05bece96292`) | SRVGGNetCompact, num_feat 64, num_conv 32, ×4, PReLU | 2,444,910 | `3d3c04150b0ab45817d2713af38a4ea08292aab9cff9650a95e51dbc36babdc5` |
| `realesrgan-x4plus-anime-6b-fp16.onnx` | [RealESRGAN_x4plus_anime_6B.pth](https://github.com/xinntao/Real-ESRGAN/releases/download/v0.2.2.4/RealESRGAN_x4plus_anime_6B.pth) (release v0.2.2.4, 17,938,799 bytes, sha256 `f872d837d3c90ed2e05227bed711af5671a6fd1c9f7d7e91c911a61f155e99da`) | RRDBNet, num_block 6, num_feat 64, num_grow_ch 32, ×4 | 9,029,798 | `848d16d1941a6ebd60e8df47b74669d85461ad26328e546097e0ac351133569c` |

The `.pth` sizes were checked against the GitHub release asset sizes (the releases do not publish digests;
the `.pth` sha256 values above were computed on download, 2026-10-07). The worker checks the `.onnx`
sha256 before using or caching a file (`lib/image-upscale.ts`, `UPSCALE_MODELS`).

## Conversion (2026-10-07)

Script: `scripts/realesrgan-to-onnx.py` (run next to the two `.pth` files; needs torch, onnx, onnxconverter-common).

1. Architectures re-implemented from `realesrgan/archs/srvgg_arch.py` (Real-ESRGAN) and
   `basicsr/archs/rrdbnet_arch.py` (BasicSR, Apache-2.0); weights loaded with `strict=True`
   (`params` for general-x4v3, `params_ema` for anime_6B).
2. `torch.onnx.export` (PyTorch 2.14, opset 17, TorchScript exporter) with input `input` [1,3,h,w] and
   output `output` [1,3,4h,4w], dynamic h/w. Input is RGB in 0–1; output is clamped to 0–1 by the app.
3. `onnxconverter_common.float16.convert_float_to_float16(keep_io_types=True)` — weights stored as fp16,
   inputs/outputs stay float32. ONNX Runtime's CPU/WASM engine computes in float32 by inserting casts;
   WebGPU computes in fp16 (only used when the adapter has `shader-f16`).

Numerical check against PyTorch fp32 on a 96×64 crop (ONNX Runtime 1.30 CPU), absolute error in 8-bit levels:

| Model | fp32 ONNX max | fp16 ONNX max | fp16 ONNX mean |
| --- | --- | --- | --- |
| general-x4v3 | 0.001 | 1.31 | 0.041 |
| anime_6B | 0.003 | 2.78 | 0.054 |

Tiling: overlap 32 input px with linear blending over the middle half of the overlap. Against whole-image
inference on a 320×213 illustration the 99.9th-percentile error was 0.5–0.7 levels (general, tiles 192–128)
and 2.3–2.5 levels (anime_6B); 8–16 px overlap left visible seams (p99.9 up to 15 levels), so 32 is the minimum.

If a file is ever replaced, give it a new name: `/models/realesrgan/*` is served as immutable.
