This is learned super-resolution, not ordinary canvas interpolation. The Apache-2.0 Swin2SR model uses a Swin Transformer V2 architecture trained for 2x image super-resolution. It reconstructs edges and textures from local image patterns. That can look sharper than stretching pixels, but it cannot recover factual detail that was never present.
The AI runs in a dedicated browser worker through Transformers.js 4.2.0. It prefers roughly 31 MB of FP16 ONNX weights on WebGPU and falls back to roughly 21 MB of q8 weights with WebAssembly on the CPU. Runtime files download separately and browsers may cache them. The tool caps each source side at 512 pixels, source area at 262,144 pixels, output area at 1,048,576 pixels, and estimated source/output tensor buffers at 20 MB.