Thank you, Mateo. I am using a lightweight custom orchestration layer written in Python instead of heavy frameworks. It gives me precise control over context window updates and message history formatting, which reduces token overhead significantly.