Hobbyist offloads model prefill to an iPhone, 44% faster on a memory-tight Mac
An external iPhone can partially bypass the prefill bottleneck of running large models on low-memory Macs; the numbers are one user's self-test, generation speed unchanged.
ImportanceLocalEvidenceE2 unreplicatedWrite-upQuick
A Reddit user used an open-source tool to move part of a large model's layers onto an iPhone 17 Pro Max, cutting prefill time for Qwen3.8-27B on a 24GB MacBook Pro: at 16K context, speed rose from 109 to 157 tokens per second, a 44% gain.
According to IT之家's October 4 report citing Wccftech, the user kept the first 40 of every 256-token batch's layers on the Mac's M4 Pro and streamed activations to the iPhone's A19 Pro GPU for layers 41 to 64; older context is compiled onto the phone's Neural Engine, cutting single-token write time from 279ms to 176ms at 140K context. Gains were 35% at 8K and 29% at 32K.
The tool, called backburner, is open source on GitHub. The limits are clear: within 64K context the phone does not speed up text generation, which stays entirely on the Mac, so the benefit applies only to long-context preprocessing.