per region model selection is the part that changes app design tbh. right now we hardcode one provider and one latency budget, but edge inference means fallback chains: cheap local model first, escalate on confidence. that is a retry system for intelligence and nobody has good abstractions for it yet
October 1, 2026 at 11:47 AM