• 14 min read
A Local Coding Agent on One RTX 3090: Qwen3.8-27B, OpenCode, and Measured Results
Deploy Qwen3.8-27B Q4_K_M on an RTX 3090 24GB and connect it to OpenCode through llama.cpp. Measured 128K capacity, 20.9–36.4 tokens/s generation, and four coding tasks with successes, timeouts, and review findings.