2 months ago
276b4d4`WORKFLOW_MAX_INFLIGHT_AGENTS = 12` suits hosted APIs. Local inference engines have far smaller request pools and per-model context budgets, where a 12-wide fan-out fails outright rather than queueing — observed as four children dying with "context size has been exceeded" at once, and separately as "worker local total request limit reached (43/32)". Adds `workflowMaxConcurrentAgents`, resolved from the config tiers and settable live through the session config patch. Clamped to [1, 16]: below 1 a workflow could not progress, and above AgentControl's MAX_ACTIVE_CHILDREN_PER_PARENT the extra slots would only produce spawn errors. Non-finite values fall back to the default so a corrupted config reads as "unset" rather than maxing out fan-out. Harness side only; the desktop settings control follows. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Kk98XjGKJrKp5ynbiNFMFu
Parentcb1ef38