FRD 0004 — Dynamic workflows¶
1. Summary¶
Add experimental Dynamic Workflows support to the markdown-first Azure Functions
Agents Runtime. Any workflow-enabled agent can ask the runtime to launch a
Durable Functions-backed DAG of tool and wait tasks, observe progress through
built-in endpoints/UI, and receive final workflow notifications in the chat
session. Workflow task tools are authored under the existing tools/ directory
but opt into Durable Activity execution explicitly with a new @workflow_tool
decorator; normal plain-function tool discovery remains backward compatible.
Workflow-enabled agents can also start the same Durable workflows from any
supported Markdown-declared trigger; the trigger starts the workflow
asynchronously and does not wait for it to finish.
The next evolution adds deterministic, data-driven control flow: a task can be
skipped by a constrained when predicate or expanded over a bounded JSON array
with for_each, while preserving Durable replay safety, owner authorization,
resource limits, deterministic fan-in, and observable node state.
Evolution¶
This FRD evolves with the experimental Dynamic Workflows surface. The initial
design assumed one workflow-enabled main.agent.md and session-only workflow
identity. PR #112 added Markdown-declared trigger starters, PR #117 added
Workflow Sub Agents, and PR #151 extends the same feature to every
workflow-enabled agent with agent/session isolation. The
multi-agent addendum
records only that extension's behavioral and architectural delta instead of
repeating the base workflow design.
The next operational evolution adds human-readable DTS dashboard labels without
changing workflow plans or registered Azure Function names. Orchestrations use
the existing workflow session agent_name plus the fixed -orchestration
suffix; tool and Sub Agent Activities use their concrete tool name or agent
slug.
2. Motivation / problem¶
Today agents can call tools directly through the Microsoft Agent Framework (MAF) during a chat turn. That works well for short, latency-sensitive work, but it is awkward for work that:
- needs multiple dependent tool calls that would otherwise require repeated model round-trips;
- can fan out independent evidence gathering in parallel;
- needs a durable wait without holding a worker or client connection open;
- produces large intermediate results that should stay out of the model context;
- should survive host restarts or a user reconnecting later.
Dynamic Workflows introduces a new authoring surface, so the first release needs
to make workflow tools easy to place, hard to register accidentally, and
consistent with the runtime's existing capability-filtering model. The agreed
model uses the existing tools/ directory as the single placement surface,
preserves normal plain-function tool discovery, and requires @workflow_tool to
explicitly opt a function into the Durable Activity execution path.
3. Goals / Non-goals¶
Goals
- Enable
workflows.enabled: truefor any agent to register Durable workflow management tools and a Durable orchestrator/activity engine. - Add
workflows.excludeso workflow filtering matches existing exclude-style capability UX (tools.exclude,mcp.exclude,skills.exclude). - Keep sample
function_app.pyminimal so workflow authoring is expressed through agent markdown plustools/. - Add
@workflow_toolas an explicit workflow authoring decorator for functions placed intools/. - Preserve existing normal
tools/behavior: public plain functions and@toolvalues continue to become normal MAF tools. - Support four clear authoring cases:
- workflow-only:
@workflow_tool; - normal-only: public plain function or
@tool; - both:
@toolplus@workflow_tool, or separate adapters sharing internal business logic; - neither:
_-prefixed helper. - Skip workflow-incompatible functions during workflow registration with a clear warning rather than failing startup when safe to do so.
- Keep discovery read-only and keep Azure Functions/Durable registration in the registration/integration stage.
- Enable every supported Markdown-declared trigger on a workflow-enabled agent to start Dynamic Workflows through the existing runner.
- Document the workflow authoring surface in
docs/workflows.md,docs/front-matter-spec.md, anddocs/architecture.md. - Add constrained conditional execution without embedding a general-purpose expression language in workflow plans.
- Add bounded runtime fan-out over JSON arrays and deterministic fan-in over the expanded results.
- Apply the existing workflow owner policy and runtime ceilings to every materialized task instance.
- Expose skipped, expanded, running, and aggregated states through the shared workflow status contract.
- Allow a plan to declare bounded
execution.retryfor tool and stateless Workflow Sub Agent tasks, then execute that policy through Durable native Activity retry without changing policy-free histories.
Non-goals
- Hand-authored workflow YAML/markdown templates; workflow plans remain
LLM-authored through
start_workflow. - Per-task concurrency and retry observability settings. Timeout and continue-on-error are included in the execution-policy extensions below.
- Sub-orchestrations, nested/stateful Sub Agent tasks, MCP Tasks integration, or cross-app workflow coordination. Stateless leaf Sub Agent tasks are in v1.
- Changing normal MAF tool execution semantics.
- Automatically promoting every compatible plain function into a workflow tool.
- General-purpose expressions, arbitrary code evaluation, loops other than bounded array iteration, or a visual workflow designer.
- Configurable resource ceilings and large-result offload; those are tracked by planning issue #1279.
4. Proposed design¶
| Pipeline stage | Module(s) | Change |
|---|---|---|
| discover | discovery/tools.py, _function_tool.py |
Load tools/*.py once, preserving normal FunctionTool discovery while also discovering explicit workflow tool declarations. Add a public workflow_tool decorator that records workflow metadata without making the function a normal MAF tool by itself. |
| translate | config/schema.py, config/merge.py, registration/capabilities.py |
Parse and validate the public workflow config shape (enabled, optional exclude, and independent subagents) and compute concrete capabilities without hard-coding the v1 workflow agent. Unknown workflow excludes warn, mirroring tools.exclude. |
| register | app.py, workflows/integration.py, workflows/registry.py, workflows/engine.py, registration/endpoints.py, registration/triggers.py |
The app composition root freezes one immutable policy per workflow-enabled agent, registers one app-wide Durable blueprint and complete execution catalogs, then threads the matching policy and Durable client through each agent's endpoints and declared triggers. |
| execute | workflows/tools.py, workflows/engine.py, workflows/activity.py, workflows/native_retry.py, workflows/context.py, runner.py, registration/_handlers.py, public/index.html |
MAF invokes workflow management tools (start_workflow, status/list/cancel/terminate). Runtime validation uses the same agent policy that generated prompt guidance. Durable Activities reauthorize against the currently deployed policy before invoking registered workflow tools or fresh stateless leaf specialists. A persisted plan-authored execution policy selects native Durable retry and its result envelope without consulting current declarations. UI polls workflow status and injects terminal notifications. |
Task execution policy¶
Planning issue #1278 is delivered as independent vertical slices. The first slice adds plan-authored retry without changing discovery or registration:
{
"id": "reserve_inventory",
"type": "tool",
"tool": "reserve_inventory",
"args": {"order_id": "order-42"},
"execution": {
"retry": {
"max_attempts": 4,
"backoff": {
"initial": "PT1S",
"multiplier": 2,
"max": "PT10S"
}
}
}
}
WorkflowTaskExecution is validated at submission and translated once into an
EffectiveWorkflowTaskExecution wire value persisted with each tool or Sub
Agent Activity input. The effective policy contains the exact Durable retry
mapping. At Activity execution, the stable idempotency key is derived from the
persisted workflow and node-instance identities rather than stored in that
policy. Replay always uses call_activity, conditionally supplying its
retry_policy argument and selecting the expected result envelope only from
the persisted input. Histories without the new execution payload retain the
legacy call and result shape.
Durable RetryPolicy has no exception predicate. Policy-aware Activities
therefore classify failures at the worker boundary, derive retryability from a
closed runtime-owned mapping, raise a private versioned marker for retryable
failures, and return a sanitized structured outcome for terminal failures. Tool
authors opt transient or terminal application failures into that mapping with
WorkflowRetryableError and WorkflowTerminalError. A stateless Workflow Sub
Agent timeout is transient; other provider, model, and configuration exceptions
remain terminal because the runtime cannot safely infer that replaying them will
succeed. Exhaustion decoding recognizes only the runtime's private marker and
surfaces its bounded application error code without leaking exception text.
The retry schedule is bounded at submission. max_attempts, backoff.initial,
backoff.multiplier, and backoff.max have fixed limits, and the ceiling of
all native retry delays must fit a one-hour admission cap. Durable's finite
retry_timeout remains unset because the SDK evaluates it against real
wall-clock time while replaying history. Per-attempt deadlines are added by the
timeout slice below and share that same admission cap.
Attempt timeout and continuation¶
Use execution to set retry, attempt timeout, and failure continuation for tool
and Sub Agent tasks. Tools can also declare retry and timeout through
@workflow_tool(retry=..., timeout=...).
{
"id": "reserve_inventory",
"type": "tool",
"tool": "reserve_inventory",
"args": {"order_id": "order-42"},
"execution": {
"timeout": "PT30S",
"retry": {
"max_attempts": 4,
"backoff": {"initial": "PT1S", "multiplier": 2, "max": "PT10S"}
},
"continue_on_error": true
}
}
Fields and defaults. Each field is optional. Defaults apply after tool declarations are combined with the plan.
| Field | Where to set it | If absent from both locations | Rules |
|---|---|---|---|
retry |
Plan or tool decorator | One attempt; no retry delay | max_attempts is 1-5, including the first attempt. More than one attempt requires backoff. The tool's retry policy replaces the plan's retry policy. |
timeout |
Plan or tool decorator | No runtime attempt deadline | ISO 8601 duration from PT1S to PT10M. The tool's timeout replaces the plan's timeout. Host and specialist limits still apply. |
continue_on_error |
Plan only | false; a task failure fails the workflow |
Boolean. true permits only the failure kinds listed below. |
Configuration patterns. R means a valid retry policy, such as the one in
the example. Tool and plan settings combine independently for each field.
| Tool declaration | Plan execution |
Effective behavior |
|---|---|---|
| None | Omitted | No execution payload. Keep existing behavior. |
| None | {"timeout":"PT30S"} |
One attempt with a 30-second deadline. |
| None | {"retry":R} |
Use R; no runtime attempt deadline. |
| None | {"continue_on_error":true} |
One attempt. Continue after a permitted failure. |
| None | {"continue_on_error":false} |
One attempt. Keep existing failure timing. |
| None | {"timeout":"PT30S","retry":R,"continue_on_error":true} |
Apply the deadline to each attempt. Continue after a terminal failure or after retries are exhausted. |
timeout="PT20S" |
Omitted | One attempt with the tool's 20-second deadline. |
retry=R |
{"timeout":"PT30S"} |
Use the tool's retry policy and the plan's timeout. |
timeout="PT20S" |
{"retry":R,"timeout":"PT30S"} |
Use the plan's retry policy and the tool's 20-second timeout. |
retry=R |
A different retry policy | Use the tool's retry policy. |
| Any | {} or null |
Reject the plan. |
Validation and persistence. Retry errors keep
workflow_retry_policy_invalid and workflow_retry_schedule_exceeded.
Explicit execution: null keeps workflow_retry_policy_invalid.
Use workflow_execution_policy_invalid for an empty execution object, invalid
timeout or continuation values, and a new decorator timeout policy applied to a
non-tool task. Existing decorator retry errors keep their current code.
For a policy without retry, use the existing
WorkflowRetryPolicy(max_attempts=1) conversion to persist max_attempts and
durable_retry_policy. Do not add a separate wire format.
The persisted timeout_ms and continue_on_error keys are NotRequired.
Write them only when declared in the effective policy. A task that uses neither
keeps the payload written by the retry-only runtime. Payload compatibility alone
does not ensure replay compatibility: preserve the existing scheduler path as
specified under Failure timing.
Admission and platform limits. The one-hour admission cap is a library
validation rule, not an Azure execution guarantee. For each materialized task
instance, the sum of retry delays and max_attempts * timeout must not exceed
one hour. Without a timeout, only retry delays count toward this cap.
Queue delays, host downtime, and other execution overhead are not included.
The cap does not limit the total elapsed time of an instance, wave, or workflow.
The host's functionTimeout setting in host.json also limits Activity
execution. The following values are from the
Azure Functions hosting limits,
checked on 2026-09-14. They describe host limits, not Python version availability
on each plan.
| Hosting plan | Default functionTimeout |
Maximum | Conditions |
|---|---|---|---|
| Consumption | 5 minutes | 10 minutes | Legacy plan. |
| Flex Consumption | 30 minutes | No fixed maximum | Scale-in grace period: 60 minutes. Platform-update grace period: 10 minutes. |
| Premium | 30 minutes | No fixed maximum | Scale-in grace period: 60 minutes. Platform-update grace period: 10 minutes. |
| Dedicated | 30 minutes | No fixed maximum | Always On is required for unbounded execution. Platform-update grace period: 10 minutes. |
| Container Apps | Normally 30 minutes | No fixed maximum | With zero minimum replicas, the default depends on the triggers. |
For a finite functionTimeout, set execution.timeout lower and allow time for
Activity setup, cleanup, and result return. For example, a 10-minute attempt
deadline cannot provide a 10-minute wait on a host with a 5-minute timeout.
A host timeout can restart the Python worker before it returns
workflow_task_timeout. Do not classify a host failure as an attempt deadline
without the runtime's failure result. No fixed maximum does not guarantee
uninterrupted execution.
Host-failure diagnostics. The Functions host detects its execution timeout, but this library cannot always identify that cause from a Durable failure. The Activity worker can stop before it records a diagnostic.
In the timeout PR, add a warning through the shared logger when the orchestrator receives an Activity failure without a valid runtime failure classification. Keep the original failure and all retry, continuation, and cancellation rules. Do not emit this warning for classified handler failures, successful results, or unrelated scheduler errors. Suppress the warning during history replay.
Include the workflow ID and, when available, the failed task instance ID and persisted attempt timeout. Do not guess the failed task from wave order or log task arguments or raw exception content. Use a message such as:
The workflow received an unclassified Activity failure. A host timeout or worker restart may have interrupted the Activity. Check Azure Functions host logs and Application Insights, if enabled. Compare the effective functionTimeout setting with the task timeout.
The diagnostic guidance must include host.json and application-setting
overrides, such as AzureFunctionsJobHost__functionTimeout. State a host timeout
as a confirmed cause only when supported failure data identifies it. Do not
infer it from elapsed time or an exception message. No diagnostic is guaranteed
if the orchestrator does not receive the failure, or if redelivery succeeds.
This warning is part of the timeout PR, not the later telemetry work.
The HTTP response limit of 230 seconds is separate. A workflow starter returns without waiting for the workflow to finish; this HTTP limit is not a deadline for the whole workflow. Do not detect the hosting SKU or change plan validation from deployment settings in this extension.
Attempt deadline. The Activity applies the persisted deadline around the
handler invocation. An expired deadline uses the existing handler_transient
kind with error_code: workflow_task_timeout. Do not add a timeout kind.
Durable retries under the persisted policy only if attempts remain.
The error code identifies the attempt deadline without a new failure kind.
The deadline limits the wait, not the lifetime of the underlying work.
A synchronous tool handler runs through asyncio.to_thread; it can keep running
after the timeout. An asynchronous handler receives cancellation, but external
work that it started can still complete. A retry can thus overlap earlier work.
Handlers must use the stable idempotency_key to prevent duplicate effects;
the key alone does not prevent them. A tool handler that raises TimeoutError
itself remains execution_unknown. Only expiration of the runtime's deadline
produces workflow_task_timeout.
A Workflow Sub Agent already carries its own resolved specialist timeout, which
may be tighter than the attempt deadline. The two bounds are independent and the
first to expire wins. Both use handler_transient. The specialist timeout keeps
its existing subagent_timeout error code; the outer attempt deadline uses
workflow_task_timeout. Neither is derived from the other, so an authored
execution.timeout never widens or narrows a specialist's own bound.
Continuation. A task whose persisted policy set continue_on_error commits a
sanitized failure result instead of failing the workflow, so its dependents run.
Durable retries a retryable failure first, if attempts remain. Only a terminal
failure or a failure after the last attempt can be continued. Continuation also
works after one failed attempt without a retry or timeout setting.
Continuability is derived from the persisted failure kind through a
runtime-owned closed map rather than added to the outcome envelope, whose exact
key set must keep validating failures written by the previous runtime:
| Failure kind | Continuable | Why |
|---|---|---|
handler_transient |
yes | No attempts remain; this includes attempt and specialist timeouts |
handler_terminal |
yes | A handler-declared application failure |
execution_unknown |
yes | An unexpected handler exception is still the handler's |
authorization |
no | Continuation must not route around the authorization boundary |
handler_contract |
no | A malformed history or handler is not an application outcome |
An opaque Durable failure carries no classification at all and therefore fails the workflow: the runtime cannot show it to be a continuable application failure.
A continued node reuses the existing completed state and commits exactly the
bounded object
— the same controlled-failure shape the orchestrator already returns, with no
aggregate results snapshot embedded, so a for_each expansion whose instances
all fail cannot grow quadratically. A downstream ${node.result...} reference or
when predicate therefore has one stable shape to read, and no scheduler state
or persisted aggregate contract changes.
Completion and errors. If every task fails with a permitted failure and
continuation enabled, the workflow still completes. Completed means that the
control flow finished, not that every task succeeded.
Keep Durable's Completed and the scheduler's completed states.
The later status/UI slice must distinguish clean completion from
Completed with errors. Base that distinction on failures that the runtime
actually continued, not a search for failed keys in user results.
Do not add a scheduler state or a new runtime_status value in these two PRs.
Until the status/UI slice ships, the current UI can show ordinary completion;
users must inspect the task results to see continued failures.
Failure timing. A wave is a group of task instances that the scheduler waits for together. The current scheduler raises Durable exceptions during that wait. It processes returned terminal failure outcomes only after the wave completes. These are different failure paths; do not change both to immediate failure.
If no instance in a wave has persisted continue_on_error: true, keep the
existing wait and result-processing path. This includes old histories,
timeout-only policies, and explicit continue_on_error: false. Preserve the
failure cause, result-application order, and cancellation order.
Only a wave with at least one persisted continue_on_error: true uses the new
per-instance processing path. In that path, record a failure and keep waiting
only if the failed instance permits continuation and its failure kind permits
it. Raise all other failures immediately, without waiting for unrelated tasks or
timers. Share this decision logic between the static and dynamic schedulers.
Select the path from persisted input, not current tool declarations.
If cancellation is selected, restore the wave and discard its recorded outcomes,
as in the current cancellation path.
Downgrade. Use of existing failure kinds does not make new policies safe on
an older runtime. The older runtime ignores timeout_ms and
continue_on_error. It stops enforcing deadlines and fails nodes that should
continue. Let in-flight workflows finish, or terminate them, before a runtime
downgrade.
Delivery plan¶
Deliver timeout and continuation as two stacked implementation PRs. The timeout
PR makes retry optional and includes decorator timeout support. The continuation
PR adds continue_on_error to the same policy format. Each PR includes tests,
sample coverage, and user documentation. Telemetry and status changes are
separate later work.
| Slice | Purpose | Scope | Dependencies | Compatibility | Review focus |
|---|---|---|---|---|---|
| Execution foundation (PR #193) | Make plan-authored execution.retry usable for tool and stateless Workflow Sub Agent tasks |
Retry schema and bounds; persisted effective policy; failure classification; Activity envelope; idempotency context; Durable mapping; static and dynamic dispatch; exhaustion sanitization; tests, docs, sample, and real-host E2E | Durable Functions Python 2.x migration PR #189 (merged) | Tasks without persisted execution data retain legacy dispatch and {"id","result"} history; no decorator or catalog surface is introduced |
Persisted-input replay selection, failure trust boundary, bounded schedules, Sub Agent timeout behavior |
@workflow_tool integration (PR #207) |
Allow tool declarations to provide retry and define decorator-over-plan precedence | Decorator metadata, discovery, registry/catalog propagation, submission-time precedence, focused tests and docs | Execution foundation (PR #193, merged) | Additive at submission; reuses the same persisted effective policy and does not change replay | Metadata propagation and precedence |
| Attempt timeout (implementation PR 1) | Limit the wait for each attempt | execution.timeout; @workflow_tool(timeout=...); per-field precedence; optional retry with a one-attempt default; persisted deadline; handler_transient with workflow_task_timeout; tests, docs, and sample |
PRs #193 and #207 (merged); base main |
No scheduler timing change; absent timeout keeps existing behavior | Policy propagation, deadline behavior, retry limits, replay |
| Continuation (implementation PR 2) | Obtain results after selected task failures | Plan-only execution.continue_on_error; bounded failure results; shared continuation logic in both schedulers; tests, docs, and sample |
Stacked on the attempt-timeout PR | Only waves with persisted continuation enabled use the new path | Failure kinds, cancellation order, old histories, for_each |
| Observability and status | Expose retry/timeout lifecycle telemetry and structured status | Telemetry, status contract, UI/docs; distinguish clean completion from completion with continued failures | Earlier execution-policy slices | Additive status version | Stable external lifecycle vocabulary; runtime-owned error tracking |
Authoring / API surface¶
Frontmatter¶
Workflow enablement remains explicit on each participating agent:
---
name: Incident Triage Assistant
description: Investigates incidents by gathering evidence in parallel.
builtin_endpoints: true
workflows:
enabled: true
exclude:
- expensive_diagnostic_tool
---
workflows.enabled:bool;trueenables Dynamic Workflows for that agent.workflows.exclude: optionallist[str]; filters discovered workflow tool names out of the effective workflow tool set.- Durable backend and task hub configuration stay in
host.jsonand app settings, not frontmatter. - No separate workflow-agent, role, or starter field is required. Invocation remains controlled independently by the agent's trigger and built-in endpoints.
Markdown-declared trigger starters¶
When a supported Markdown-declared trigger belongs to a workflow-enabled agent, registration adds a Durable client input to that generated Function. The handler passes the bound client, workflow enablement, the agent identity slug, and trigger-specific system guidance to the existing runner. Workflow-disabled handlers retain their original signatures.
start_workflow schedules the orchestration and returns a workflow_id to the
agent. The initial trigger Function ends after that agent turn instead of
polling for terminal workflow status. An HTTP caller receives the immediate
agent response; non-HTTP triggers have no response channel, so applications can
provide a workflow tool that delivers the eventual result to an appropriate
destination. This evolution adds no new frontmatter fields.
Tool decorators¶
Normal tool behavior stays unchanged:
The public plain function above remains a normal MAF tool only. It does not become a workflow Activity target.
Workflow-only tools opt in with @workflow_tool. The decorator attaches
workflow metadata and returns the original callable/object so it does not make a
function a normal MAF tool by itself:
from azure_functions_agents import workflow_tool
@workflow_tool(description="Fetch recent log lines for a service.")
def fetch_logs(args: dict[str, object]) -> dict[str, object]:
service = str(args["service"])
return {"service": service, "errors": 12}
Both direct MAF tools and workflow tools can be expressed by applying both
decorators when the callable contract is intentionally shared. Decorator order
should not affect discovery: @workflow_tool attaches metadata to a plain
callable or to a FunctionTool, and discovery also checks the wrapped
FunctionTool.func for workflow metadata.
from azure_functions_agents import tool, workflow_tool
@tool
@workflow_tool(description="Get current service health.")
def get_service_health(args: dict[str, object]) -> dict[str, object]:
return {"service": args["service"], "status": "healthy"}
The reverse order is also valid:
@workflow_tool(description="Get current service health.")
@tool
def get_service_health(args: dict[str, object]) -> dict[str, object]:
return {"service": args["service"], "status": "healthy"}
The single-callable "both" pattern supports synchronous and async callables that
satisfy both the MAF and workflow Activity contracts. With @tool(schema=Params),
the runtime converts the workflow argument dictionary to the Pydantic model
before it calls the handler. Both decorator orders are supported.
When normal tools use a Pydantic model but workflow Activities use dict
arguments, authors can also share internal business logic and expose separate
adapters when the input or output contracts differ:
from pydantic import BaseModel
from azure_functions_agents import tool, workflow_tool
class HealthParams(BaseModel):
service: str
def _get_health(service: str) -> dict[str, object]:
return {"service": service, "status": "healthy"}
@tool
def get_service_health(params: HealthParams) -> str:
return str(_get_health(params.service))
@workflow_tool(name="get_service_health")
def get_service_health_workflow(args: dict[str, object]) -> dict[str, object]:
return _get_health(str(args["service"]))
Helpers remain _-prefixed:
def _require_service(args: dict[str, object]) -> str:
service = args.get("service")
if not isinstance(service, str) or not service:
raise ValueError("service is required")
return service
Workflow tool execution contract¶
For v1, a workflow tool handler must:
- be synchronous or
async(the Activity awaitsasynchandlers; see #139); - accept one
dict[str, Any]argument; - return a JSON-serializable value;
- avoid relying on chat-turn-local runtime state;
- be appropriate for Durable Activity execution, including background and parallel execution.
The runtime should warn and skip functions that are clearly incompatible, such
as declaration-only tools, reserved names, duplicate names, or
handlers whose signature cannot accept the workflow dict argument.
Reserved workflow tool names are the workflow management tools injected by the
runtime: start_workflow, get_workflow_status, list_workflows,
cancel_workflow, and terminate_workflow.
Duplicate detection is scoped to the workflow registry only. It is valid for a normal MAF tool and a workflow tool to share the same name intentionally; that is the expected shape for tools that support both direct chat use and workflow DAG execution.
Compatibility¶
- Existing normal tools remain backward compatible:
- public plain functions continue to be auto-wrapped as normal
FunctionToolinstances; - existing
@toolusage remains a normal MAF tool. @workflow_toolalone must not accidentally enter the normal plain-function fallback path.- Sample
function_app.pystays minimal; samples use the sametools/plus@workflow_toolauthoring model expected of users. @workflow_toolaccepts only supported v1 metadata (name,description,public) until retry/timeout metadata is implemented. Unknown keyword arguments fail fast at startup so authors do not think unsupported policy knobs are active.
Workflow Sub Agents¶
[!IMPORTANT] This extension is part of the Dynamic Workflows v1 surface. PR #151 applies the same contract to every workflow-enabled agent. The
samples/workflow-subagents-preview/directory is the runnable single-agent Workflow Sub Agent sample.
The extension lets a workflow-enabled agent authorize existing Markdown agents as DAG nodes:
---
name: Support Coordinator
workflows:
enabled: true
subagents:
- agent: pr_status_analyst
when: Review one pull request and summarize its current status
- agent: actionable_report_writer
when: Combine pull-request summaries into an actionable portfolio report
---
workflows.subagents and the top-level chat-time subagents: list are
independent capability grants. Both are deny-by-default when omitted.
workflows.subagents may reference a specialist used only by Workflows.
Unknown, duplicate, and self references fail during app composition. As with a
top-level subagents: reference, an authorized Workflow-only specialist does
not need its own trigger or built-in endpoint. when is the routing hint shown
to the coordinator's plan-authoring model; when omitted, the specialist's
description is used. The subagents items are translated into typed
configuration during app composition rather than re-parsed by registration or
execution code.
The static grant and every runtime plan are enforced independently. Before a
plan starts, each sub_agent.agent must be present in the owning agent's
workflows.subagents grant. An unauthorized or unknown slug rejects the plan;
the Activity also fails closed if its catalog lookup cannot resolve the
already-authorized slug. The immutable agent-specific policy used for prompt
guidance is the same policy used for plan validation. Composition constructs one
independent immutable policy per workflow-enabled agent without changing the
node or Activity contract.
The Workflow plan uses a sub_agent task:
{
"id": "analyze_pr_42",
"type": "sub_agent",
"agent": "pr_status_analyst",
"task": "Review pull request https://github.com/owner/repo/pull/42 and summarize its current status."
}
The reduce node uses the same task type and depends on every map result:
{
"id": "write_report",
"type": "sub_agent",
"agent": "actionable_report_writer",
"task": "Create an actionable report from PR 42: ${analyze_pr_42.result.text}; PR 43: ${analyze_pr_43.result.text}.",
"depends_on": ["analyze_pr_42", "analyze_pr_43"]
}
task must be a self-contained string and may template upstream results. A
successful v1 node returns
{"agent": "pr_status_analyst", "text": "..."};
downstream tasks can reference ${analyze_pr_42.result.text}. Independent Sub
Agent tasks can fan out without dependencies, and another authorized Sub Agent
can depend on all of them to reduce their summaries. Status and lineage remain
owned by the parent Workflow and identify the execution by parent Workflow id,
node id, and specialist slug. Leaf-only means that the specialist cannot start
another Workflow or delegate again. A Sub Agent Activity is not an independently
queryable workflow instance: built-in status surfaces report it only as a parent
node, including the currently scheduled node ids while a wave is running.
The specialist runs as itself with a fresh context and its own instructions,
model, timeout, normal tools, MCP servers, skills, and web_request setting. It
does not inherit the parent's tools or conversation history. In v1 it also
receives no request-scoped sandbox, Workflow management tools, or delegate_*
tools.
The specialist's configured timeout is enforced inside the async Agent Activity
around Agent.run(task). The Functions host's activity/function timeout remains
an outer limit, so the observable upper bound is the shorter of the specialist
timeout and the host limit. A timeout raises from the Activity and fails the
parent Workflow; it is never returned as a success-shaped result.
| Concern | Proposed v1 | Deferred to v2 |
|---|---|---|
| Execution | One stateless Agent Activity per leaf node; no child orchestration | Stateful or bounded multi-level execution |
| Result | Fixed {agent, text} envelope |
response_schema-validated output |
| Failure | Activity failure or timeout fails the parent Workflow | Retry and continue-on-error policy |
| Retry | No automatic retry; use the specialist's timeout | Idempotent retry with attempts/backoff |
| Cancellation | Parent stops scheduling; an already-dispatched model call is best-effort | Stronger activity interruption where supported |
| Context | Self-contained task only |
Explicit context-sharing policy, if justified |
The v1 runtime does not configure automatic Durable retries. The task and result authoring contract should remain unchanged if a runtime-managed Durable retry policy is added later. Before enabling it, the implementation must define idempotency, retryable failure kinds, maximum attempts/backoff, and how repeated model or tool side effects are surfaced.
Even without configured retry options, Durable Activity delivery is at-least-once. A worker failure can therefore repeat a model call or specialist tool side effect. v1 does not claim exactly-once Agent execution: specialist tools used from a Workflow should tolerate re-execution, and terminal publishers should use stable destination identities or equivalent idempotent writes. The PR-status sample overwrites the request's specified Blob path so repeated publication converges on the same report instead of creating duplicate outputs.
Reviewer note: positive capability allowlists¶
Today specialist tools, skills, and mcp capabilities inherit the
project-wide inventory and can only be narrowed with exclude (or disabled
entirely). The proposal preserves that existing behavior, but durable background
execution makes the lack of a positive allowlist a least-privilege concern:
adding a new project capability can make it available to existing specialists
without editing their definitions.
A future capability proposal could add an explicit form such as:
This syntax is illustrative only and is not accepted as part of the Workflow Sub Agent contract in this draft. Review should decide whether positive allowlists are a prerequisite, a parallel feature, or a later hardening step.
Multi-agent workflow isolation addendum (PR #151)¶
This addendum supersedes the original main.agent.md-only assumption. It does
not introduce new frontmatter or DAG syntax: every agent with
workflows.enabled: true receives the existing workflow tools and may start
workflows through whichever triggers or built-in endpoints it independently
exposes.
The implementation calls such an agent a workflow-enabled agent. Its
workflow_agent_slug defines the authorization namespace but is not an
authoring keyword.
App-wide execution and per-agent authorization¶
One Function App registers one Durable orchestrator, one copy of each Activity,
one complete workflow-handler catalog, and the existing immutable
AgentCatalog. Registering that engine once prevents duplicate Azure Functions
when several agents enable workflows. Complete catalogs answer what exists; they
do not grant access.
Composition separately freezes one WorkflowPlanPolicy per workflow-enabled
agent. That policy contains only the agent's workflow tools after
workflows.exclude and its deny-by-default workflows.subagents grants. The
same value drives prompt guidance, start-time plan validation, and
defense-in-depth Activity authorization. One agent's exclusions cannot
unregister a handler another agent may use.
Agent and session isolation¶
ResolvedAgent.slug is the stable agent identity on chat, MCP, HTTP-trigger, and
non-HTTP-trigger paths. Workflow management is scoped by
(workflow_agent_slug, session_id) internally. Durable instance IDs begin with a
32-hex-character (128-bit) truncated SHA-256 digest over an unambiguous
length-delimited encoding of both values, followed by the existing random UUID
suffix. Raw slugs and session IDs are not exposed in instance IDs.
Active-workflow limits, list, status, cancel, terminate, and HTTP polling all require both components. A mismatched agent or session returns the same not-found/empty result as an unknown workflow, so two agents remain isolated even when a caller deliberately reuses one session ID.
This intentionally changes the experimental workflow-ID prefix from the legacy session-only 48-bit digest. New application tools and routes do not manage pre-upgrade IDs. Operators must drain or terminate legacy instances through Durable Functions or DTS tooling before upgrading.
Activity-time reauthorization¶
Capability-bearing Activities carry workflow_agent_slug and check the currently
deployed policy immediately before shared-catalog dispatch:
- tool Activities require the task tool in
policy.allowed_tools; - Workflow Sub Agent Activities require the specialist in
policy.allowed_subagents; and - a missing policy, handler, or Agent catalog entry fails closed with a non-sensitive error and correlated logs.
Persisting the start-time policy as indefinitely authoritative would defeat revocation. Removing a workflow-enabled agent while another remains, or tightening its grants, therefore makes a pending disallowed Activity fail rather than continue with stale authorization. Durable orchestrator replay performs no mutable policy lookup.
Final-agent removal lifecycle¶
The last workflow-enabled agent is a special deployment edge case. With no agent
policy, normal composition intentionally returns to a plain FunctionApp to
avoid Durable overhead in apps that do not use workflows. A plain
FunctionApp, however, has no registered workflow orchestrator or Activities.
Consider an instance whose next Activity has been scheduled but has not yet executed:
sequenceDiagram
participant H as Task Hub
participant A as Deployment with final workflow agent
participant P as Plain FunctionApp after agent removal
H->>A: Activity work item is pending
Note over A,P: Final workflow-enabled agent is removed
H--xP: No Activity Function is registered to receive the work item
Note over H: Instance remains non-terminal instead of reaching policy rejection
This differs from ordinary policy revocation: the work item cannot reach
require_workflow_agent_policy() and fail because the Function that executes
that check is absent. The application does not expose a drain-mode environment
variable for this edge case. Publishing such a switch would create a durable
customer compatibility commitment without resolving privileged direct Durable
starts or establishing whether lifecycle ownership belongs in this runtime or
Durable itself.
Before removing the final workflow-enabled agent, operators should stop new starters and use Task Hub tooling to let existing instances finish or terminate them. The supported long-term behavior is deferred to issue #161, which tracks runtime/Durable ownership and a remediation that does not prematurely add public surface.
Runnable proof¶
samples/per-agent-workflows/ contains two non-main workflow-enabled agents with
different workflow-tool exclusions and Workflow Sub Agent grants. Repository E2E
automation starts both with the same session ID against Azure Storage and DTS,
proves each reaches a terminal state using only its own capabilities, and checks
that cross-agent status access returns 404.
Data-driven control flow (Issue #1276; in review)¶
The current workflow contract is an arbitrary but static DAG: every task id and
dependency edge exists when start_workflow validates the plan. Static roots can
already fan out and a later task can fan in through depends_on, but the model
must enumerate every item before submission. That prevents a workflow from
adapting to a bounded collection returned by a tool or Sub Agent and forces
irrelevant branches to run even when an upstream result makes them unnecessary.
This extension keeps the LLM-authored DAG as the control plane and adds two optional fields to each existing task type:
when: a constrained predicate that decides whether the logical task or materialized task instance runs.for_each: a full-value reference to an upstream JSON array. The runtime materializes one instance of the task per array element.
No frontmatter field is added. Existing plans that omit both fields retain their
current validation, scheduling, result, and status behavior. The optional fields
are omitted with exclude-unset/exclude-none serialization when absent so static
plan model dumps and Durable wire payloads do not gain null fields.
Before and after¶
The diagram contrasts the static fan-out/fan-in already supported before this extension with the data-driven flow proposed here. Blue nodes are existing capabilities; green and amber nodes are new in Issue #1276.
flowchart TB
subgraph BEFORE["Before Issue #1276 — static DAG (already supported)"]
direction LR
B0["LLM authors every task<br/>and every dependency"]:::existing
B1["analyze_pr_a"]:::existing
B2["analyze_pr_b"]:::existing
B3["analyze_pr_c"]:::existing
B4["summarize<br/>fixed fan-in"]:::existing
B0 --> B1
B0 --> B2
B0 --> B3
B1 --> B4
B2 --> B4
B3 --> B4
end
subgraph AFTER["With Issue #1276 — data-driven DAG"]
direction LR
A0["LLM authors logical tasks only<br/>discover → analyze → summarize"]:::existing
A1["discover result<br/>[PR A, PR B, PR C]"]:::existing
A2["Runtime resolves for_each<br/>checks owner policy + node budget"]:::new
A3["analyze[0]<br/>when = true → run"]:::new
A4["analyze[1]<br/>when = false → skipped"]:::skipped
A5["analyze[2]<br/>when = true → run"]:::new
A6["analyze logical result<br/>ordered [0, 1, 2] aggregate"]:::new
A7["summarize<br/>consumes ${analyze.result}"]:::existing
A0 --> A1 --> A2
A2 --> A3
A2 --> A4
A2 --> A5
A3 --> A6
A4 --> A6
A5 --> A6
A6 --> A7
end
classDef existing fill:#dbeafe,stroke:#2563eb,color:#172554
classDef new fill:#dcfce7,stroke:#16a34a,color:#052e16
classDef skipped fill:#fef3c7,stroke:#d97706,color:#451a03
| Before this extension | Added by Issue #1276 |
|---|---|
| The LLM enumerates every concrete task id before submission. | The LLM authors one logical for_each task; the runtime creates bounded [index] instances. |
Parallel roots and fixed depends_on fan-in are supported. |
Fan-out size comes from an upstream JSON array at runtime. |
| Every ready task runs. | Constrained when predicates can skip a logical task or individual instance. |
| Downstream templates reference separately authored task results. | The logical task exposes one source-ordered aggregate, including explicit skipped positions. |
Pipeline mapping¶
| Pipeline stage | Module(s) | Change |
|---|---|---|
| discover | No change | Dynamic control flow does not discover new application files or capabilities. The data-driven authoring skill is a runtime-owned packaged asset, not an app-discovered skill. |
| translate | workflows/schema.py, workflows/tools.py |
Extend the agent-facing and runtime plan schemas with typed when and for_each fields. Validate syntax, upstream references, static tool/Sub Agent targets, and logical DAG structure before scheduling. |
| register | app.py, registration/capabilities.py, workflows/integration.py, runner.py, pyproject.toml |
Give direct workflow-enabled agents the packaged data-driven-workflows skill independently of project-skill filtering, reduce the runtime-owned addendum to a short on-demand load instruction, and retain local field guidance in the start_workflow tool schema. Build a direct-role capability copy after catalog construction so delegated roles keep the original project-only paths. Durable blueprint registration and owner policy construction remain unchanged. |
| execute | workflows/engine.py, workflows/tools.py, public/index.html |
Decode the validated Durable JSON payload into typed task/state boundaries, then deterministically resolve collections and predicates, materialize bounded instances, schedule them under the existing parallelism cap, aggregate results in source order, publish structured progress, and normalize controlled failures into stable envelopes. |
Progressive authoring guidance and typed execution¶
Detailed data-driven authoring guidance uses MAF skill progressive disclosure
instead of occupying every workflow-enabled turn's system instructions. The
runtime packages a data-driven-workflows SKILL.md. Its name and short
description are visible to direct workflow-enabled agents; the body is loaded
only when the model decides a plan needs runtime conditions or collection fan-out.
The narrow skill description is the selection pointer; the shared addendum does
not mention the skill because E2E evaluation showed that even a qualified
addendum pointer encouraged speculative loads for fixed DAGs. Static workflows
therefore do not pay the detailed control-flow context cost.
This runtime-owned skill is part of workflows.enabled, not the application's
discovered skills/ inventory. It remains available when project skills are
disabled or filtered and is not advertised in discovery summaries. After the
two-pass catalog has retained each agent's project-only AgentCapabilities,
registration creates a shallow direct-role capability copy with the built-in
path appended. Trigger and built-in endpoint handlers receive that copy; catalog
entries and delegated-role construction retain the original capabilities, so
the same agent does not inherit the skill when it runs as a leaf Sub Agent
without workflow management tools.
data-driven-workflows is a runtime-reserved skill name. Application
composition fails with a clear error if project discovery returns that name
rather than relying on MAF's order-dependent skill de-duplication. Public docs
remain the human contract; the packaged skill is the model-readable contract,
and tests keep their core grammar aligned.
The asset lives at
src/azure_functions_agents/workflows/skills/data-driven-workflows/SKILL.md.
workflows/integration.py resolves it relative to its own __file__, and
pyproject.toml includes workflows/skills/** as package data. A wheel test
must build and inspect or install the non-editable artifact because editable
installs do not prove package-data inclusion.
The Durable boundary still persists JSON, but scheduler internals do not operate
on open-ended dict[str, Any] values. TypedDict and Literal contracts describe
the persisted workflow payload, task variants, logical/instance states, and
materialized instances. Newly added dynamic keys are NotRequired; the Durable
payload boundary reads them with compatibility defaults once so in-flight
static instances created before this evolution replay unchanged. The
orchestrator trusts the already submission-validated persisted payload instead
of reconstructing Pydantic models during replay. Untrusted pre-Pydantic
validation input remains Mapping[str, object] and is narrowed explicitly.
The dynamic orchestrator retains its Durable generator and every yield in one
top-level function. Deterministic synchronous helpers own materialization,
iteration binding, runnable selection, result application, and cancellation
restoration through one typed state object. This limits nesting without moving
Durable side effects or replay-sensitive control flow behind opaque abstractions.
Constrained when contract¶
when is an object rather than a string expression:
{
"id": "notify",
"type": "tool",
"tool": "send_notification",
"args": {"incident": "${classify.result.incident}"},
"depends_on": ["classify"],
"when": {
"ref": "${classify.result.should_notify}",
"operator": "equals",
"value": true
}
}
The contract is intentionally small:
refmust be one full reference to an upstream result or, insidefor_each, the current${item}/${item.path}/${index}local.operatoris exactlyequalsornot_equals.valuemust be a JSON scalar (null, boolean, number, or string).- Comparison is type-sensitive JSON scalar equality. There is no coercion, truthiness, ordering, regex, boolean composition, function call, or access to environment/runtime state.
- A missing path, malformed reference, non-scalar resolved value, or unsupported operator is an error; it never silently evaluates to false.
For a normal task, the predicate is evaluated once after all dependencies
complete. For a for_each task, the collection is resolved first and the
predicate is evaluated independently for each bound item. Evaluation order is:
resolve dependencies, resolve for_each when present, bind ${item} /
${index}, evaluate when, and only for a true predicate resolve executable
args or Sub Agent task templates. A false predicate therefore does not
resolve unused executable value fields; it marks the corresponding logical task
or instance skipped, schedules no Activity/timer, and produces null for that
result position. A skipped task still satisfies downstream depends_on edges.
Skip does not propagate automatically. A descendant that should be part of the
same conditional branch must declare its own when; this keeps branch behavior
visible in the authored plan and avoids an implicit dependency-reachability
language. A full ${skipped.result} reference resolves to null. Traversing
below it, such as ${skipped.result.field}, produces the controlled
workflow_reference_unresolved failure because null has no traversable path.
Bounded for_each contract¶
for_each is available on tool and sub_agent tasks and must be one full
upstream-result reference that resolves to a JSON array:
{
"id": "analyze",
"type": "sub_agent",
"agent": "pr_status_analyst",
"task": "Analyze pull request ${item.url} at input index ${index}.",
"depends_on": ["discover"],
"for_each": "${discover.result.pull_requests}"
}
The task's target (tool, agent, or wait) remains static and is validated
against the owner's immutable policy before the workflow starts. Only value
fields (args, a Sub Agent's task, and when.ref) may use the fixed iteration
locals:
${item}returns the current element with its native JSON type.${item.path.to.field}traverses the current element using the same deterministic dictionary/list path rules as upstream result templates.${index}returns the zero-based integer index.
item and index are reserved and rejected as authored task ids in every
plan. This keeps the iteration-local namespace unambiguous without contextual
precedence rules, including for references such as ${item.result}.
Aliases, nested for_each, cross-instance references, item-dependent
depends_on, and templated tool or Sub Agent names are not supported. An array
element may be any JSON value, although a referenced item path must be valid for
that element. wait tasks may use when but cannot use for_each: repeated
identical timers add no data-driven behavior because wait deadlines cannot
reference iteration locals.
The template grammar, validation walker, and runtime resolver are extended to
recognize ${item}, ${item.path}, and ${index}. Those forms are rejected
outside a for_each task, and the existing unmatched-token defense continues to
reject every other ${...} shape.
Materialized instance ids are runtime-owned and use
<logical-task-id>[<zero-based-index>], for example analyze[0]. They are
visible in status and diagnostics but cannot appear in authored depends_on or
template references. Authored task ids continue to allow letters, numbers,
underscore, and hyphen only; [ and ] are rejected, reserving the rendered
instance-id namespace for the runtime. Materialization and scheduling always use
the numeric (logical_task_id, index) tuple as the ordering key, with logical
task id as the outer key when multiple tasks become ready together. The scheduler
must not sort the rendered instance-id strings because analyze[10] sorts before
analyze[2] lexicographically and would violate source-index wave selection even
though that string order is itself replay-deterministic. These rules make the
same persisted inputs and upstream results produce the same instance ids and
Durable scheduling history on replay.
An empty array is valid: no instances run, the logical node immediately becomes
aggregated, and its result is [].
Deterministic fan-in¶
A for_each logical node completes only after all of its materialized instances
have completed or been skipped. Its result is an array aligned with the source
collection:
[
{"index": 0, "status": "completed", "result": {"summary": "ready"}},
{"index": 1, "status": "skipped", "result": null}
]
The array is always ordered by source index, never by Activity completion order.
A downstream task depends on the logical id ("depends_on": ["analyze"]) and
can consume the complete collection with ${analyze.result} or traverse a known
position with the existing dotted/list-index syntax. It cannot depend on or
reference an individual runtime-owned instance id.
This is aggregation of already-completed instance results, not a new reducer language. Domain-specific reduction remains an ordinary workflow tool or authorized Sub Agent task.
Limits and authorization¶
The existing static plan cap still limits authored logical tasks. In addition, the runtime maintains a materialized-node budget:
- each non-iterated task consumes one node;
- each
for_eacharray element consumes one node, including an element later skipped bywhen; - an empty expansion consumes no materialized nodes;
- before scheduling any instance from an expansion, the engine rejects the
whole expansion if it would make the workflow exceed
MAX_NODES; - individual ready instances are scheduled under the existing
MAX_PARALLELISMcap.
Counting skipped instances prevents a large collection from bypassing the node limit through a predicate. Runtime-configurable ceilings remain out of scope for this extension and belong to planning issue #1279.
Every materialized instance inherits the already-validated task type and static target. Materialization re-applies the same immutable owner policy before dispatch as defense in depth; collection data can change arguments or Sub Agent instructions but cannot select a different tool or specialist. Dynamic control flow therefore does not broaden the workflow's capability grant.
Stable failures¶
Submission and runtime-controlled failures use the same flat error fields.
start_workflow preserves the current top-level "error": "<message>" field
for compatibility and adds error_code plus bounded context such as node_id
and path. Runtime-controlled failures add failed: true and partial results
to the same shape; the shared status adapter exposes that terminal output as
runtime_status: "Failed":
{
"failed": true,
"error": "Task 'analyze' for_each did not resolve to an array.",
"error_code": "workflow_iteration_not_array",
"node_id": "analyze",
"path": "${discover.result.pull_requests}",
"results": {"discover": {"pull_requests": "omitted from this example"}}
}
The failure phase and status behavior are fixed:
| Code | Submission validation | Runtime resolution | Status behavior |
|---|---|---|---|
workflow_task_id_reserved |
Authored task id is item or index |
N/A; rejected before scheduling | Submission returns the flat error directly |
workflow_condition_invalid |
Malformed predicate, unsupported operator, invalid literal/reference shape | Resolved predicate value is not a JSON scalar | Submission returns the flat error directly; runtime output maps to Failed |
workflow_reference_unresolved |
Unknown/non-upstream task, iteration local outside for_each, malformed reference |
Missing dict key, invalid/out-of-range list index, or traversal through a scalar/null |
Submission returns the flat error directly; runtime output maps to Failed |
workflow_iteration_not_array |
N/A; result type is not knowable yet | for_each resolves to a non-array JSON value |
Runtime output maps to Failed |
workflow_node_limit_exceeded |
Authored logical task count exceeds the static limit | A resolved expansion would exceed the materialized-node budget | Submission returns the flat error directly; runtime output maps to Failed |
Submission failures occur before a Durable instance is created and therefore
have no runtime_status. Runtime failures are observable through
get_workflow_status, list_workflows, and the HTTP status endpoint as a normal
status envelope whose runtime_status is Failed and whose output is the flat
failure object above. Messages may improve over time; callers key on
error_code. results contains every logical result committed before the
failure. A per-instance failure uses the runtime-owned instance id in node_id
(analyze[3]), while collection materialization and aggregation failures use the
logical id (analyze).
Provider, model, and tool failures remain governed by the existing sanitized
failure behavior and the separate reliable-execution work in issue #1278.
Runtime occurrences of the four controlled failures above are returned by the
orchestrator rather than raised. status_envelope() and _is_active_status()
map output.failed is True to runtime_status: "Failed", mirroring the existing
cooperative-cancel mapping. Existing raise-based template-resolution paths are
migrated to this single returned envelope so an unresolved runtime reference has
one stable shape whether it occurs in normal args, a Sub Agent task, when, or
for_each. Unexpected engine invariants and Activity/provider failures continue
to raise and use native Durable failure behavior. Status consumers must check
output.failed is True before interpreting output as the controlled flat
schema; other Failed instances retain the native/opaque Durable failure output.
Structured status¶
The status envelope keeps its existing top-level fields, but custom_status
becomes a versioned JSON object for dynamically controlled workflows. The legacy
free-form string is status schema version 1; structured snapshots use version 2:
{
"schema_version": 2,
"counts": {
"logical_total": 3,
"materialized_total": 4,
"completed": 2,
"skipped": 1,
"running": 1
},
"nodes": {
"discover": {"state": "completed"},
"analyze": {
"state": "running",
"expanded_count": 3,
"instances": {
"analyze[0]": {"state": "completed"},
"analyze[1]": {"state": "skipped"},
"analyze[2]": {"state": "running"}
}
}
}
}
Logical node states are pending, running, skipped, expanded,
aggregated, completed, or failed; instance states omit expanded and
aggregated. A for_each node is expanded after materialization, running
while any runnable instance is in flight, and aggregated after its ordered
result array is committed. The shared status tools and HTTP endpoint pass this
object through unchanged, and the built-in UI renders the states rather than
parsing progress text. Static v1 workflows may continue returning their current
string custom_status; clients must accept either shape during the experimental
compatibility window.
Sample¶
The Dynamic Workflow sample for this extension must demonstrate:
- a discovery tool returning a bounded JSON array;
- one
for_eachtool or Sub Agent node whose predicate skips at least one item; - a downstream task consuming the ordered aggregate via the logical node id;
- status output showing expanded, running, skipped, and aggregated states; and
- deterministic completion on both Azure Storage and DTS Durable backends.
Durable Task Scheduler display names¶
The runtime attaches the well-known durabletask.displayName tag when it starts
an orchestration and schedules a tool or Workflow Sub Agent Activity. DTS uses
that tag as the primary label in orchestration lists, sequence/flow views, and
detail panels while retaining the registered function name in metadata.
The orchestration label is derived deterministically as
<agent_name>-orchestration. Here, agent_name is the existing workflow
session value currently populated from the resolved agent slug on production
invocation paths. The runtime does not ask the LLM to invent a workflow title
and does not add a field to start_workflow.
Tool Activities use the workflow tool name. Workflow Sub Agent Activities use
the authorized agent slug. Expanded for_each instances intentionally share
the same display name; their distinct runtime task IDs remain in the Activity
input and details. Timer tasks are unchanged.
The pinned Durable Functions client exposes orchestration tags through
schedule_new_orchestration, so workflow startup moves from deprecated
start_new(..., client_input=...) to
schedule_new_orchestration(..., input=..., tags=...).
The Azure Functions one-argument compatibility orchestration context does not
expose Activity tags. The engine therefore registers its orchestrator using the
supported native two-argument Durable Task contract. The orchestration payload
becomes the second argument; native call_activity(..., input=..., tags=...)
and module-level task combinators replace compatibility-only helpers. The
engine explicitly normalizes the native replay-safe UTC timestamp before
comparing it with timezone-aware absolute wait deadlines. Activity names,
inputs, authorization, scheduling order, status payloads, workflow IDs, and
results remain unchanged.
5. Decisions log¶
| # | Decision | Options considered | Choice | Decided by | Date |
|---|---|---|---|---|---|
| 1 | Workflow execution backend | Direct chat tool loop / in-process scheduler / Durable Functions | Durable Functions orchestrator + Activity engine | Human + Agent | 2026-07-01 |
| 2 | Workflow enablement surface | Always on / agent frontmatter flag / global config only | workflows.enabled: true on main.agent.md |
Human + Agent | 2026-07-01 |
| 3 | Workflow tool placement | Dedicated workflow_tools/ / existing tools/ |
Existing tools/ directory |
Human | 2026-07-06 |
| 4 | Workflow tool opt-in | Auto-promote compatible plain functions / @tool(workflow=True) / explicit @workflow_tool |
Explicit @workflow_tool decorator |
Human | 2026-07-06 |
| 5 | Normal plain function behavior | Stop auto-wrapping / keep existing normal tool discovery | Keep existing plain-function discovery for normal MAF tools | Human | 2026-07-06 |
| 6 | Workflow filter style | exclude list / no filtering |
Use workflows.exclude to match existing capability filtering |
Human | 2026-07-06 |
| 7 | Workflow-only functions | Require duplicate wrappers / @workflow_tool only / config-only exclusion |
@workflow_tool only means workflow-only and must not become normal MAF tool |
Human + Agent | 2026-07-06 |
| 8 | Future workflow metadata | Separate config maps / decorator kwargs / postpone with no surface | Reserve @workflow_tool(...) for future retry/timeout/etc. metadata |
Human + Agent | 2026-07-06 |
| 9 | Incompatible workflow candidates | Fail all startup / silently skip / warn and skip where safe | Warn and skip incompatible workflow tool declarations where safe | Human | 2026-07-06 |
| 10 | Workflow filtering stage | Apply workflows.exclude in integration/register / compute concrete workflow tools in capabilities |
Compute the concrete workflow tool set before registration so registration consumes objects, not exclude policy | Agent | 2026-07-06 |
| 11 | Dual decorator order | Require one order / support both orders | Support both orders by attaching workflow metadata to both callables and FunctionTool objects |
Agent | 2026-07-06 |
| 12 | Record trigger support | Create a second Dynamic Workflows FRD / evolve this FRD | Update FRD 0004 because Markdown-declared trigger support extends the existing feature without redesigning it | Human | 2026-07-23 |
| 13 | Declared-trigger scope | Add named trigger types individually / use generic trigger registration | Add the Durable client binding generically to every supported Markdown-declared trigger for the workflow-enabled main agent | Human + Agent | 2026-07-17 |
| 14 | Trigger lifetime | Wait for terminal status / start asynchronously | End the initial trigger Function after the agent starts the workflow; Durable execution continues independently | Human + Agent | 2026-07-17 |
| 15 | Workflow Sub Agent authorization | Reuse the top-level list / add a mode flag / use a Workflow-owned grant | Add independent, deny-by-default workflows.subagents |
Human | 2026-07-23 |
| 16 | First execution boundary | Recursive delegation / bounded nesting / leaf-only | v1 is leaf-only; bounded multi-level execution is v2 | Human | 2026-07-23 |
| 17 | Specialist context | Copy parent state / share history / self-contained task | Run with the specialist's own static capabilities and a self-contained task only | Human | 2026-07-23 |
| 18 | Failure and retry | Recoverable result / automatic retry / fail parent without retry | Sub Agent failure fails the parent Workflow; v1 has no automatic retry | Human | 2026-07-23 |
| 19 | Successful result | Plain text / schema-dependent result / fixed envelope | Return {agent, text}; defer response_schema to v2 |
Human | 2026-07-24 |
| 20 | Sub Agent runtime boundary | Direct Activity / one child orchestrator per node / shared child orchestrator | Invoke each stateless Sub Agent directly as an Activity; retain status and lineage on the parent node | Human + Chris Gillum | 2026-07-24 |
| 21 | Dependency on per-agent Workflows (#109) | Wait for #109 / ship main-only then extend | Ship the existing main.agent.md owner scope now, while keeping engine and policy boundaries reusable by #109 |
Human | 2026-07-24 |
| 22 | Documentation audiences | Explain internals in every document / separate maintainer and customer surfaces | Keep decisions and Durable internals in the FRD/architecture; make samples and authoring docs independently understandable to customers | Human + Chris Gillum | 2026-07-24 |
| 23 | Sub Agent failure diagnostics | Expose provider errors / one generic message / bounded error code plus correlated logs | Keep provider details out of Durable history, expose a stable non-sensitive error code, and correlate detailed logs by Workflow ID, node ID, and specialist slug | Human + Laveesh Rohra | 2026-08-03 |
| 24 | Record multi-agent support | Create FRD 0009 / evolve this FRD | Keep the change as an addendum to FRD 0004 because it extends the existing experimental feature without adding a new authoring contract | Human + Laveesh Rohra | 2026-08-12 |
| 25 | Workflow agent identity | Display name / source path / endpoint-specific name / canonical slug | Use app-wide unique ResolvedAgent.slug on every channel as workflow_agent_slug inside the authorization implementation |
Agent | 2026-08-10 |
| 26 | Workflow isolation scope | Session only / agent only / agent plus session | Scope application management by (workflow_agent_slug, session_id) so equal session IDs across agents remain isolated |
Human | 2026-08-10 |
| 27 | Existing workflow IDs | Dual-format fallback / migration map / no application fallback | Accept the experimental ID change, preserve Durable/DTS operator access, and require upgrade drain guidance | Human | 2026-08-10 |
| 28 | Durable registration lifetime | Once per agent / once per app | Register the Durable engine and complete execution catalogs exactly once per app | Agent | 2026-08-10 |
| 29 | Per-agent policy | Mutable process global / request-time reconstruction / immutable slug-keyed catalog | Freeze one independent WorkflowPlanPolicy per workflow-enabled agent during composition |
Agent | 2026-08-10 |
| 30 | Activity authorization | Trust start-time validation / persist start-time policy / reauthorize deployed policy | Reauthorize tool and Sub Agent Activities against the currently deployed agent policy so restrictive changes fail closed | Human | 2026-08-10 |
| 31 | Non-HTTP trigger management | Add an app-wide index / share one synthetic session / generated invocation session | Keep generated sessions and no new application index; use Durable/DTS tooling for app-wide operations | Human | 2026-08-10 |
| 32 | Authoring schema | Add owner/config fields / reuse current workflow config | Reuse existing fields; derive identity from the canonical agent slug | Agent | 2026-08-10 |
| 33 | Ownership digest width | Keep 48 bits / store literal identity / increase digest | Use a 128-bit truncated SHA-256 prefix over length-delimited agent/session input | Human | 2026-08-10 |
| 34 | Workflow agent eligibility | Require a dedicated starter and fail composition / allow every enabled agent | Treat every agent with workflows.enabled: true as workflow-enabled; invocation surfaces remain independent. This supersedes the earlier provisional fail-composition rule. |
Human | 2026-08-11 |
| 35 | Customer sample boundary | Put sender/verifier helpers in the sample / separate internal automation | Keep the sample documentation-led and directly runnable; keep E2E automation under tests/scripts |
Human | 2026-08-11 |
| 36 | Final-agent removal | Always register Durable / documentation-only drain / explicit runtime retention | Add opt-in drain mode that blocks application starts and retains Durable registration until Task Hub tooling confirms no non-terminal instances; ordinary non-workflow apps remain plain FunctionApp |
Human | 2026-08-12 |
| 37 | Exported compatibility helpers | Remove production-dead helpers / retain shared state / isolate compatibility state | Retain exported registry and one-shot integration helpers without an unrelated breaking change, but keep their registration token out of production WorkflowSessionContext and never authorize production execution from the singleton fallback |
Agent | 2026-08-11 |
| 38 | Trigger decorator resolution | Add a shared resolver / duplicate capability validation / retain registration-local fallback | Keep the registration-local connector_trigger to generic_trigger fallback and avoid an unrelated hard failure for non-workflow agents |
Agent | 2026-08-11 |
| 39 | Activity-wave failure propagation | Rely on Durable wrapper behavior / explicitly rethrow the failed wave result | Explicitly rethrow failed task_all results so policy denials retain their actionable error instead of degrading to a secondary TypeError |
Agent | 2026-08-11 |
| 40 | Internal workflow-agent terminology | owner_slug / agent_slug / workflow_agent_slug |
Use workflow_agent_slug throughout workflow plumbing and persisted payloads: it identifies the top-level agent that starts, authorizes, and namespaces the workflow without colliding conceptually with a delegated Sub Agent |
Human | 2026-08-12 |
| 41 | Final-agent lifecycle public surface | Keep AZURE_FUNCTIONS_AGENTS_WORKFLOW_DRAIN_MODE / remove it and track the lifecycle gap / always register Durable |
Remove the environment variable before release and track the edge case in #161; publishing an operational switch would create a customer compatibility commitment before runtime versus Durable ownership is resolved. This supersedes Decision #36. | Human + Laveesh Rohra | 2026-08-13 |
| 42 | Record dynamic control flow | Create a separate FRD / evolve FRD 0004 | Evolve FRD 0004 because conditions and iteration extend the existing workflow plan and engine contract | Human | 2026-08-13 |
| 43 | Condition surface | General expression string / JSON predicate object / boolean-only reference | Use a constrained JSON predicate with scalar equals / not_equals; reject missing paths and type mismatches |
Agent | 2026-08-13 |
| 44 | Iteration surface | Embedded loop expression / for_each full array reference / generated child plan |
Use one for_each upstream-array reference with fixed ${item} and ${index} locals |
Agent | 2026-08-13 |
| 45 | Dynamic instance identity | Value hash / random id / source index | Derive runtime-only <logical-id>[<index>] ids from source order |
Agent | 2026-08-13 |
| 46 | Fan-in result | Completion-order list / keyed object / source-aligned envelopes | Aggregate source-ordered {index, status, result} envelopes under the logical node id |
Agent | 2026-08-13 |
| 47 | Skipped dependency behavior | Auto-propagate / block descendants / explicit descendant conditions | Do not auto-propagate; satisfy dependencies with null, require each conditional descendant to declare when, and fail controlled dotted traversal below null |
Agent | 2026-08-13 |
| 48 | Dynamic resource accounting | Count only executed Activities / count every materialized item / separate unlimited expansion | Count every materialized item, including skipped items, against MAX_NODES; retain MAX_PARALLELISM |
Agent | 2026-08-13 |
| 49 | Dynamic status contract | Continue free-form strings / event log / versioned structured snapshot | Add a versioned custom_status object while accepting legacy strings for static plans |
Agent | 2026-08-13 |
| 50 | Controlled error compatibility | Replace the error shape / messages only / stable code alongside existing shape | Preserve the existing error message field and add stable codes plus bounded context | Agent | 2026-08-13 |
| 51 | Iterated wait tasks | Permit identical timers / template deadlines / reject iteration | Reject for_each on wait; keep when available for conditional waits |
Agent | 2026-08-13 |
| 52 | Controlled runtime failure provenance | Raise native Durable failure / return envelope and status-map / Activity wrapper | Return one stable envelope and map output.failed to Failed; reserve native raises for unexpected and Activity failures |
Agent | 2026-08-13 |
| 53 | Dynamic instance namespace | Permit all authored ids / escape collisions / reserve bracket suffixes | Restrict authored ids to letters, numbers, underscore, and hyphen; reserve [index] suffixes for runtime instances |
Agent | 2026-08-13 |
| 54 | Dynamic control-flow design approval | Revise individual Decisions 43-53 / approve the proposed set | Approve Decisions 43-53 as proposed and advance to implementation | Human (TsuyoshiUshio) | 2026-08-14 |
| 55 | Iteration-local task ids | Permit shadowing with local precedence / reject only ambiguous references / reserve iteration-local names | Reserve item and index as task ids for every plan so ${item} and ${index} always denote for_each iteration locals |
Human (TsuyoshiUshio) | 2026-08-19 |
| 56 | Data-driven authoring context | Keep the full shared addendum / point only to public docs / package an on-demand MAF skill | Package data-driven-workflows; keep only a short load instruction in the shared addendum so detailed grammar enters context only for data-driven plans |
Human (TsuyoshiUshio) | 2026-08-19 |
| 57 | Built-in skill scope | Treat it as a filterable project skill / make it a workflow system capability / expose it to all agents | Give it only to direct workflow-enabled agents as a runtime system capability, independent of skills: false; do not expose it to leaf Sub Agent execution or application skill discovery |
Human (TsuyoshiUshio) | 2026-08-19 |
| 58 | Durable scheduler typing | Continue dict[str, Any] / reconstruct Pydantic models during replay / typed JSON wire contracts |
Use TypedDict/Literal contracts and one typed mutable scheduler state while retaining JSON persistence and avoiding Pydantic reconstruction inside the orchestrator | Human (TsuyoshiUshio) | 2026-08-19 |
| 59 | Dynamic scheduler structure | Keep one deeply nested function / class-based orchestrator / shallow deterministic phase helpers | Keep the generator/yield boundary in _run_dynamic_workflow and extract synchronous typed helpers for materialization, selection, result application, and cancellation restoration |
Human (TsuyoshiUshio) | 2026-08-19 |
| 60 | Review-follow-up design approval | Defer to a later PR / implement individual nits / approve Decisions 56-59 together | Approve the progressive-disclosure and typed phase-refactor design for this PR, subject to an independent architecture review before implementation | Human (TsuyoshiUshio) | 2026-08-19 |
| 61 | Direct versus delegated skill plumbing | Add the built-in path to shared catalog capabilities / add a second persistent capability field / create a direct-role copy after catalog construction | Keep catalog capabilities project-only and pass a shallow copy with the built-in path only to trigger and built-in endpoint registration | Agent, architecture review | 2026-08-19 |
| 62 | Built-in skill name collision | Let MAF keep the first duplicate / namespace without reservation / reserve and fail composition | Reserve data-driven-workflows and fail application composition when a project skill uses the same name |
Agent, architecture review | 2026-08-19 |
| 63 | Built-in skill packaging | Generate content in code / external public file / packaged workflow asset | Store SKILL.md under workflows/skills/data-driven-workflows, resolve relative to integration.py, add workflows/skills/** package data, and verify a built wheel |
Agent, architecture review | 2026-08-19 |
| 64 | Legacy Durable payload decoding | Require all new keys / revalidate via Pydantic / optional typed keys with one compatibility boundary | Mark dynamic keys NotRequired, apply defaults once at the persisted JSON boundary, and trust the previously validated payload internally |
Agent, architecture review | 2026-08-19 |
| 65 | Progressive-disclosure selection pointer | Keep a qualified shared-addendum pointer / rely on public docs / use only MAF skill metadata | Keep the narrow load condition in the Skill description and remove the shared-addendum pointer after E2E showed that mentioning the Skill there caused fixed DAGs to load it speculatively | Agent, E2E evidence | 2026-08-19 |
| 66 | Record DTS display names | Create a new FRD / evolve FRD 0004 | Evolve FRD 0004 because display tags are a small operational improvement to the existing Dynamic Workflows engine with no new authoring contract | Human (TsuyoshiUshio) | 2026-09-03 |
| 67 | Orchestration display label | LLM-authored title / agent display name / existing workflow session agent_name |
Use <agent_name>-orchestration so labels are deterministic and require no LLM-facing schema change |
Human (TsuyoshiUshio) | 2026-09-03 |
| 68 | Activity display label | Runtime Activity name / task id / execution target | Use the workflow tool name for tool Activities and the authorized agent slug for Workflow Sub Agent Activities | Human + Agent | 2026-09-03 |
| 69 | Durable Activity-tag integration | Access compatibility-context internals / wait for wrapper support / native two-argument orchestrator | Use the public native Durable Task orchestrator contract exposed by the pinned package; preserve existing execution semantics and cover the context migration in regression tests | Human + Agent | 2026-09-03 |
| 70 | Task execution policy delivery | Ship timeout, retry, continue-on-error, and status/observability together / decompose into stacked changes | Land Durable native retry first as an independently mergeable change, then stack per-attempt timeout with continue-on-error, then retry/timeout observability and the structured status contract | Human (TsuyoshiUshio) | 2026-08-29 |
| 71 | Retry driver | Orchestrator-managed retry timers / Durable native Activity retry / both with a selector | Use Durable native Activity retry only. An orchestrator-managed loop would duplicate scheduling Durable already owns and need its own timers, attempt state, and status vocabulary | Human (TsuyoshiUshio) | 2026-08-29 |
| 72 | Durable SDK dependency | Keep azure-functions-durable 1.x / adopt 2.x |
Depend on the Durable Functions Python 2.x migration already delivered by PR #189; 1.x cannot express the authored exponential backoff contract | Agent, architecture review | 2026-08-29 |
| 73 | Retryable-versus-terminal signaling | Retry every Activity exception / a retry predicate / raise for retryable and return for terminal | Durable RetryPolicy has no exception predicate, so the Activity raises a private versioned marker for retryable failures and returns a structured outcome for terminal ones. Retryability is derived from the failure classification, never trusted from the worker payload |
Agent, architecture review | 2026-08-29 |
| 74 | Per-attempt task timeout | Required by native retry / independent follow-up | Not required for this slice: the retry delay schedule is bounded separately from an individual attempt. An in-flight attempt remains bounded by the host functionTimeout; authored execution.timeout follows separately |
Agent, architecture review | 2026-09-02 |
| 75 | Replay of pre-retry histories | Re-resolve policy at replay from current declarations / dispatch from persisted orchestration input only | Select the retry driver and result envelope from persisted orchestration input alone, so a deployment cannot change how an in-flight workflow replays | Agent, architecture review | 2026-09-02 |
| 76 | Persisted policy forward compatibility | Strict extra="forbid" on the replayed shape / ignore unknown keys |
Ignore unknown keys on every model read back from Durable history, and require keys added by a later runtime to be optional, so histories written on either side of an upgrade still validate | Agent, architecture review | 2026-08-29 |
| 77 | Attempt number in task context | Expose the current attempt / expose only a stable idempotency key | Expose only the idempotency key. Durable owns the attempt budget and a replayed orchestration cannot observe the attempt, so publishing one would be a value handlers could not trust | Agent | 2026-08-29 |
| 78 | Retry authoring surface in the first slice | Ship decorator and plan together / plan-authored first / decorator only | Ship plan-authored execution.retry only. It is complete through validate, persist, dispatch, and exhaust without discovery or registration changes; retain @workflow_tool(retry=...) and decorator-over-plan precedence for the rebased remainder of PR #185 |
Human (TsuyoshiUshio) | 2026-09-02 |
| 79 | Workflow Sub Agent retry classification | Retry every leaf failure / reject Sub Agent retry / retry only a closed transient set | Treat a leaf TimeoutError as transient and retryable; classify all other leaf exceptions as terminal unless a future reviewed mapping proves they are safe to replay |
Agent, architecture review | 2026-09-02 |
| 80 | Retry schedule time bound | Set Durable retry_timeout / validate an authored delay-sum cap only |
Validate the one-hour delay-sum cap before start and leave Durable retry_timeout unset. The SDK compares that timeout to real wall-clock time while replaying old failure events, so a finite value can change historical scheduling after enough time passes |
Agent, final review | 2026-09-02 |
| 81 | Rebase strategy for decorator retry | Rebase the full PR #185 branch / port only the approved residual slice | Port only @workflow_tool(retry=...) metadata propagation and submission precedence onto current main; rebasing the stale full branch would reintroduce already-merged foundation changes and enlarge review scope |
Human + Agent | 2026-09-08 |
| 82 | Rebase strategy for timeout and continuation | Rebase the stale PR #186 branch / reimplement the slice on current main |
Reimplement on current main. The merged foundation (PRs #189, #193, #207) replaced the branch's versions of schema.py, activity.py, and engine.py, so every file the branch touches conflicts wholesale; a fresh slice is smaller and reviewable against the shipped contracts |
Human (TsuyoshiUshio) | 2026-09-10 |
| 83 | Shape of an execution object with no fields |
Keep retry required / allow an empty object / require at least one declared field |
Require at least one field. All three fields become optional so a timeout may be declared alone, but an empty execution would move a policy-free task onto the structured envelope for no authored reason. Rejected at submission with workflow_execution_policy_invalid |
Agent | 2026-09-10 |
| 84 | Authoring location of continue_on_error |
Task and tool declaration / task only | Task only. Whether a workflow may proceed past a failed node is a property of the plan, not of the tool; a tool cannot know which plan can tolerate its absence. timeout stays declarable on both, keeping decorator-over-plan precedence identical to retry |
Agent | 2026-09-10 |
| 85 | What the attempt deadline bounds | Cancel the worker / bound the awaited attempt | Bound the awaited attempt. A workflow tool handler runs synchronously on a worker thread through asyncio.to_thread and cannot be cancelled from outside, so a timed-out attempt is reported while it may still run; a Sub Agent delivery is asynchronous and is cancelled, though provider-side work already dispatched may still complete. Both are the at-least-once exposure the stable idempotency_key already exists for; it is documented rather than hidden |
Agent, architecture review | 2026-09-10 |
| 86 | Classifying an expired deadline | Reuse handler_transient / add a timeout kind |
Add a timeout kind, retryable. Reusing handler_transient would make an infrastructure deadline indistinguishable from a handler-declared transient failure in the failure contract the next slice reports on. A handler that raises TimeoutError itself stays execution_unknown: only the scope that actually expired reports a timeout |
Agent | 2026-09-10 |
| 87 | Carrying continuability | Add a field to the outcome envelope / derive it from the persisted kind |
Derive it. The envelope is validated by an exact key set, so adding a field would reject every failure written by the previous runtime. A runtime-owned map keeps the decision on the runtime side, where authorization and handler_contract can never be continued past |
Agent | 2026-09-10 |
| 88 | Failure timing on a wave | Fail fast always / collect every outcome when continuation is declared / decide per completed instance | Decide per completed instance. Collecting a whole wave would make a non-continuable failure wait behind a slow sibling or a long timer merely because another node opted in, and could report a different node as the cause. The selection loop instead records a failure and keeps awaiting only when that instance declared continuation and its classification is continuable; every other failure raises immediately as it does today. Both inputs are persisted or derived from the outcome, so the choice is replay-stable | Agent, architecture review | 2026-09-10 |
| 89 | Continued node representation | New failed_continued scheduler state / reuse completed with the controlled-failure result |
Reuse completed. _aggregate_dynamic_node copies instance state into the persisted {index, status, result} aggregate, so a new state would change a shipped result contract and every terminal-state predicate. The node commits the {"failed": true, ...} envelope the orchestrator already returns, giving downstream references one stable shape. Reporting the distinction is the observability slice's job |
Agent | 2026-09-10 |
| 90 | Bounded-execution ceiling with deadlines | Bound delays only / bound deadlines and delays together | Bound them together, per materialized task instance. max_attempts * timeout plus the retry delay ceiling must fit the existing one-hour admission cap, so an authored deadline cannot extend the worst-case schedule past the bound the retry slice established. The cap is deliberately not a whole-workflow bound: MAX_NODES instances scheduled MAX_PARALLELISM at a time can still exceed an hour in aggregate |
Agent, architecture review | 2026-09-10 |
| 91 | Persisting the two new keys | Always write them / write only when authored | Write only when authored. Both are NotRequired, so a task that declares neither freezes a payload byte-identical to the one the retry-only runtime wrote and replays through exactly the same path — the property that keeps an in-flight upgrade safe |
Agent | 2026-09-10 |
| 92 | Runtime downgrade with new histories in flight | Version the persisted policy for backward readability / declare downgrade unsupported | Declare it unsupported and drain or terminate in-flight workflows first, matching the existing rollback stance. An older runtime's closed retryability map rejects the timeout kind as a contract failure and ignores both new persisted keys, so a downgrade silently stops enforcing deadlines and fails continued nodes. Encoding the new kind inside an old-readable shape would mean lying about the classification |
Agent, architecture review | 2026-09-10 |
| 93 | Continuability of each failure kind | Continue any non-authorization failure / enumerate a closed map | Enumerate a closed map: timeout, handler_transient, handler_terminal, and execution_unknown are continuable; authorization and handler_contract never are, because continuation must not route around the authorization boundary or treat a malformed history as an application outcome. An unmapped kind is not continuable by default |
Agent, architecture review | 2026-09-10 |
| 94 | Content of a continued node result | Reuse the workflow failure envelope / commit a bounded object | Commit the bounded {failed, error_code, error, kind} object only. The workflow-level failure envelope carries the aggregate results map, so reusing it would embed a growing snapshot in every continued instance and let a fully failed for_each expansion grow quadratically |
Agent, architecture review | 2026-09-10 |
| 95 | Workflow Sub Agent timeout precedence | Derive one bound from the other / keep both independent | Keep both independent; the first to expire wins. The specialist's own resolved timeout keeps the handler_transient classification Decision 79 established and only the outer attempt deadline reports timeout, so an authored execution.timeout never widens or narrows a specialist bound and no shipped classification changes |
Agent, architecture review | 2026-09-10 |
| 96 | Slice sizing for timeout plus continuation | Stack timeout and continuation as two PRs / deliver one slice | Deliver one slice. Both change the same optionality of execution and the same persisted-key compatibility argument, so splitting would review that identical schema and replay change twice, and continuation is only meaningful once an attempt budget can actually be spent. Reviewability is preserved by keeping schema, Activity, and scheduler changes in separate commits |
Human (TsuyoshiUshio), Agent | 2026-09-10 |
| 97 | Requirement review and PR boundaries | Keep one implementation PR / remove decorator timeout / use two stacked PRs | REDUCE. Keep decorator timeout in scope. Deliver timeout, including decorator support, in PR 1; deliver continuation in PR 2 on top of it. Each PR includes tests, sample coverage, and docs. Telemetry and status changes remain separate. This replaces Decision 96: continuation is useful after one failed attempt without a timeout or retry setting | Human (TsuyoshiUshio), Agent | 2026-09-14 |
| 98 | Timeout failure representation | New failure kind / existing kind with a separate error code | Use handler_transient with workflow_task_timeout for the attempt deadline. Keep subagent_timeout for the specialist deadline. This replaces Decision 86 and the new-kind parts of Decisions 92, 93, and 95. Downgrade still requires workflows to finish or be terminated because older code ignores the new policy keys |
Human (TsuyoshiUshio), Agent | 2026-09-14 |
| 99 | Failure timing and replay | Process all failures immediately / preserve the old path unless continuation is enabled | Replace Decision 88 and clarify Decision 91. Current code raises Durable exceptions during the wave wait but processes returned terminal failures after the wave. Use new per-instance processing only when a wave has persisted continue_on_error: true. Otherwise preserve failure cause, result-application order, and cancellation order. Identical payloads alone do not prove replay compatibility |
Human (TsuyoshiUshio), Agent | 2026-09-14 |
| 100 | Retry omitted from an execution policy | New persisted format / existing one-attempt policy | After decorator precedence, default an absent retry policy to WorkflowRetryPolicy(max_attempts=1). Reuse the existing conversion to persist required retry fields with no delay. Apply this to timeout-only and continuation-only policies; tasks with no settings keep no execution payload |
Human (TsuyoshiUshio), Agent | 2026-09-14 |
| 101 | Host-failure guidance | Documentation only / diagnostic warning / automatic SKU validation | Add a replay-suppressed warning when the orchestrator receives an unclassified Activity failure. Give possible causes and host-log/configuration checks without changing the failure. Keep SKU detection out of scope | Human (TsuyoshiUshio), Agent; review by Victoria Hall | 2026-09-15 |
| 102 | Completion after continued failures | New scheduler state / separate result display | Keep existing execution states. Require the later status/UI slice to distinguish completion with continued failures, including when all tasks fail. Use runtime-owned continuation records, not user result keys. Document the interim UI limitation | Human (TsuyoshiUshio), Agent; review by Victoria Hall | 2026-09-15 |
| 103 | Timeout sample delivery | Add a separate timeout sample / extend the retry sample | Extend workflow-retry-policy with a carrier task whose decorator timeout overrides a longer plan timeout while the plan retry remains active. This gives one runnable app for retry success, timeout retry, and timeout exhaustion without duplicate Functions setup |
Agent | 2026-09-16 |
| 104 | Final review delivery | Keep two stacked review PRs / combine completed slices in one final review PR | Combine the completed timeout and continuation slices in one final review PR. The implementation stayed in two commits during development and passed separate implementation and testing reviews. This replaces the review boundary in Decision 97; the product scope and separate status/UI slice do not change | Human (TsuyoshiUshio), Agent | 2026-09-16 |
6. Test plan¶
- [ ] Unit:
tests/test_discovery_tools.py - plain public functions still become normal tools;
@toolvalues still become normal tools;@workflow_tool-only functions do not become normal tools;- modules can expose multiple workflow tools;
_-prefixed helpers are ignored.- [ ] Unit: dual-decorator behavior
@toolover@workflow_toolis both a normal tool and a workflow tool;@workflow_toolover@toolis both a normal tool and a workflow tool;- duplicate names are rejected only within the workflow registry, not across normal and workflow tool inventories.
- [ ] Unit: workflow discovery/registry tests
- compatible
@workflow_toolhandlers register automatically; - async handlers are accepted; incompatible handlers are skipped with warning logs;
- duplicate/reserved names are handled with clear warnings/errors;
@workflow_toolusing a reserved runtime management name such asstart_workflowis rejected;- effective workflow tool set respects
workflows.exclude. - [ ] Unit:
tests/test_workflow_integration_validation.py workflows.excludeshape validation;- unknown workflow keys fail with actionable messages.
- [ ] Unit:
tests/test_app_routes.py - workflow-enabled app startup discovers sample workflow tools from
tools/; - workflow addendum lists discovered non-excluded workflow tools.
- [ ] Fixture scenario:
tests/fixtures/config_scenarios/<next>_dynamic_workflow_tools/ tools/contains normal-only, workflow-only, both, and helper functions.- [ ] Sample tests: update
tests/test_incident_tools.pyfor the decorator-based sample layout. - [x] E2E:
tests/endtoend/test_workflow_tools_e2e.pyruns theworkflow-incident-triagesample underfunc startwith Azurite/Durable storage and a live Foundry model. A chat prompt makes the agent write the plan and callstart_workflow. The test then confirms that the workflow completes and that the Durable Activities ran the syncfetch_logsandfetch_metricshandlers, awaited the asyncfetch_deployshandler, and returned their results. - [x] Evolution #112: workflow-enabled HTTP and non-HTTP handlers receive the Durable client and trigger addendum while workflow-disabled handlers keep their existing signatures.
- [x] Evolution #112: timer and queue samples index their trigger, Durable client, orchestrator, and Activity bindings and complete model-backed local runs.
- [x] Evolution #117: Workflow Sub Agents
- validate the independent, deny-by-default
workflows.subagentsgrant; - reject a runtime
sub_agentnode whose slug is not authorized by that grant, and fail closed on an impossible catalog miss; - validate
sub_agentnode shape, authorization, DAG templates, and results; - execute map nodes as parallel Agent Activities and reduce their
{agent, text}results; - verify specialist capability isolation, timeout, failure, and cancellation;
- make
samples/workflow-subagents-preview/runnable and execute it end to end through Queue, Durable execution, fake PR tools, HTML reduction, and Blob publication, including convergence on the same Blob after repeated publication. - [x] Evolution #151: multi-agent workflows
- compose every workflow-enabled agent with an independent immutable policy;
- register one app-wide Durable engine and complete execution catalogs;
- isolate IDs and management by agent plus session and return non-existence semantics for cross-agent access;
- reauthorize capability-bearing Activities against the deployed policy;
- keep an ordinary app with no workflow-enabled agents on plain
FunctionApp; - document and track the unresolved final-agent lifecycle edge case without exposing a drain-mode environment variable;
- treat legacy session-only IDs as not-found without deleting or mutating their Durable instances;
- prove independent agents and same-session isolation against Azure Storage and DTS.
- [ ] Evolution #1276: schema and validation
- accept optional
whenon every task type andfor_eachon tool/Sub Agent tasks; - reject unsupported operators, malformed/local references outside iteration, non-upstream references, templated targets, nested iteration, and iterated waits;
- reject authored task ids outside letters, numbers, underscore, and hyphen so
they cannot collide with runtime
[index]instance ids; - reject
itemandindexas authored task ids so iteration-local references cannot be shadowed; - preserve unchanged model dumps and wire payloads for static v1 plans.
- [ ] Evolution #1276: progressive authoring guidance
- workflow-enabled direct agents always receive the packaged
data-driven-workflowsskill, including when project skills are disabled; - workflow-disabled agents and leaf Sub Agent execution do not receive it;
- the Skill description selects
when/for_eachauthoring without a shared addendum pointer, and fixed-DAG E2E does not load the Skill; - packaged-wheel tests prove the built-in
SKILL.mdis included and readable; - project discovery rejects the reserved
data-driven-workflowsname before MAF skill-provider construction; - skill and public-doc contract tests cover iteration locals, reserved ids, skip/null behavior, strict predicates, and ordered aggregation.
- [ ] Evolution #1276: typed scheduler structure
- mypy checks task variants, payload fields, state literals, and instance fields
without scheduler-local
dict[str, Any]; - malformed pre-validation input is narrowed from
Mapping[str, object]; - helper extraction preserves Durable replay ordering, yield boundaries, controlled failure envelopes, cancellation restoration, and static scheduling.
- [ ] Evolution #1276: deterministic execution
- replay produces identical instance ids, ordering, scheduling waves, skip decisions, and aggregate results;
- numeric scheduling order remains source-aligned across index 9/10 and later parallelism waves;
- empty, singleton, duplicate-value, mixed-type, and maximum-size arrays behave deterministically;
- skip does not propagate, full skipped-result references resolve to
null, and dotted traversal below a skipped result fails with a stable code; whenis evaluated before executable args/task templates, so invalid unused fields on a skipped instance are not resolved;- collection/type/path failures produce one returned controlled-failure
envelope and status-map to
Failed. - [ ] Evolution #1276: stable failure phases
- submission validation and runtime-controlled failures use the same flat
error/error_code/ bounded-context fields; - each stable code is exercised in every applicable phase, and no Durable instance is created for submission failures;
- runtime-controlled failures are returned, mapped to
Failed, and exposed unchanged by tool and HTTP status surfaces; - per-instance failures report the runtime instance id, preserve completed logical results, and remain distinguishable from opaque native Durable failures.
- [ ] Evolution #1276: limits and authorization
- expansion is rejected atomically before dispatch when the materialized node
budget would exceed
MAX_NODES; - skipped instances count against the node budget and runnable instances obey
MAX_PARALLELISM; - every expanded tool and Sub Agent instance reuses the immutable owner policy and cannot template its target.
- [ ] Evolution #1276: status and UI
- status snapshots expose skipped, expanded, running, aggregated, and failed nodes/instances;
- tools, HTTP status routes, and the built-in UI accept both legacy string and
versioned object
custom_statusvalues. - [ ] Evolution #1276: sample/E2E
- a sample discovers a collection, dynamically fans out, skips one item, and aggregates results;
- the scenario completes with deterministic output on Azure Storage and DTS.
- [ ] Evolution #DTS display names:
tests/test_workflow_registry.pyverifiesschedule_new_orchestration(..., input=..., tags=...)and the deterministic<agent_name>-orchestrationlabel;tests/test_workflow_engine.pydrives the native two-argument orchestrator contract and verifies tool/Sub Agent tags on static, dynamic, and expanded Activities;- cancellation, timers, custom status, authorization, deterministic ordering, and absolute/relative wait deadlines retain their existing behavior;
- emulator and Azure DTS runs show display tags in place of shared registered function names.
- [ ] Evolution #1278 slice 1: plan-authored retry contract
- validate the bounded retry schema and reject unsupported task types or schedules whose delay ceiling exceeds the internal retry window;
- persist the effective policy at submission and derive the stable idempotency key at Activity execution from persisted workflow and node-instance identities;
- ignore later optional keys while rejecting inconsistent persisted mappings.
- [ ] Evolution #1278 slice 1: execution and compatibility
- dispatch static and dynamic tool/Sub Agent tasks through Durable native retry only when persisted execution data selects it;
- classify tool-declared transient/terminal failures and leaf timeouts without trusting worker-supplied retryability;
- sanitize terminal and exhausted failures while retaining bounded application error codes;
- preserve policy-free dispatch, result envelopes, and replay behavior.
- [ ] Evolution #1278 slice 1: sample/E2E
- a runnable sample instructs its agent to author
execution.retryfor an idempotent workflow tool; - a real Functions host proves retry-to-completion and sanitized exhaustion against the local Durable backend.
- [x] Timeout PR: execution contract
- accept
execution.timeoutalone, reject anexecutionthat declares no field, and reject a duration outsidePT1S-PT10M; - apply
@workflow_tool(timeout=...)over a plan-authored timeout; - resolve decorator precedence per field, covering decorator-timeout with plan-retry and decorator-retry with plan-timeout;
- keep shipped retry error codes on retry shape, schedule, and explicit-
nullexecution failures while the new cases reportworkflow_execution_policy_invalid; - reject a schedule whose attempt deadlines plus retry delay ceiling exceed the one-hour admission cap;
- persist a timeout-only policy as one attempt with no retry delay when neither the plan nor the decorator declares retry;
- omit the timeout key unless declared, preserving retry-only payloads.
- [x] Timeout PR: attempt deadline
- an expired deadline uses
handler_transientwithworkflow_task_timeout; tool and Sub Agent deliveries retry only while attempts remain; - a handler that raises
TimeoutErroritself staysexecution_unknown; - a Sub Agent whose own specialist timeout is tighter than the attempt deadline
reports
subagent_timeout, and the reverse ordering reportsworkflow_task_timeout; both keep thehandler_transientkind; - a synchronous handler can finish after its deadline; asynchronous handlers receive cancellation, and external cancellation is not reported as a timeout;
- a history with no persisted deadline keeps its unbounded attempt;
- a persisted deadline outside its validated domain is a contract failure;
- deadline and retry-delay validation does not change scheduler failure timing.
- [x] Timeout PR: host-failure diagnostics
- an unclassified Activity failure emits the guidance warning with available workflow/task identity and persisted timeout, without changing the failure;
- classified handler failures, successful results, and scheduler errors do not emit the warning;
- replay suppresses the warning, and logs contain no task arguments or raw exception content;
- missing failure details or task identity do not produce a guessed cause.
- [x] Continuation PR: execution contract
- accept continuation without timeout or retry, using the existing one-attempt persisted policy when neither plan nor decorator declares retry;
- keep
continue_on_errorplan-only and reject it on the decorator; - omit the continuation key unless declared and preserve earlier payloads;
- reject an empty execution object and invalid continuation values.
- [x] Continuation PR: DAG continuation
- a continued node commits the bounded
{"failed": true, ...}object with no aggregate results snapshot, stayscompleted, and lets dependents,whenpredicates, and skip propagation run in both the static and dynamic schedulers; - an exhausted native-retry failure is continued only after the budget is spent;
- each continuable kind continues and each non-continuable kind fails, including
an opaque Durable failure and an explicit
continue_on_error: false; - a
for_eachexpansion continues with some instances failed and with all instances failed, and a wave whose instances all fail continuably still advances; - a non-continuable failure in a wave that also contains a declared-continuation instance and a pending timer still raises immediately;
- cancellation racing a continuable failure restores the wave and discards the recorded failure;
- a wave with no enabled continuation keeps immediate Durable exceptions and deferred processing of returned terminal outcomes, including authorization failures; preserve its failure cause and result-application order;
- old histories, timeout-only policies, and explicit false continuation keep cancellation order when a terminal outcome arrives before a slow sibling;
- exercise the same rules in both schedulers.
- [ ] Status/UI slice: completion with errors
- distinguish clean completion, some continued failures, and all tasks failed with continuation enabled;
- do not treat a successful user result containing
failed: trueas a runtime-continued failure. - [x] Timeout PR: sample/E2E
- include a runnable timeout example with decorator precedence;
- use a real Functions host to prove retry and exhaustion after an attempt deadline against the local Durable backend;
- configure a host timeout shorter than the attempt deadline and record the failure data available to the orchestrator. Check diagnostic guidance when it receives an unclassified Activity failure. Do not assume that the worker can log before restart.
- [x] Continuation PR: sample/E2E
- include a runnable example of a failed optional task and its dependents;
- use a real Functions host to prove continuation after timeout exhaustion, continuation after a single terminal failure, and cancellation order;
- replay old histories with a terminal failure and a pending sibling or timer.
Mock tasks alone do not reproduce Durable 2.x
when_anyand.resultexception behavior.
7. Docs impact¶
- [ ]
docs/architecture.md— add workflows to the data flow, module map, and pipeline-stage descriptions. - [ ]
docs/front-matter-spec.md— documentworkflows.enabledandworkflows.exclude. - [ ]
docs/workflows.md— document@workflow_toolauthoring and auto-registration fromtools/. - [ ]
README.md— ensure experimental workflows mention points to the sample and docs. - [ ]
samples/workflow-incident-triage/README.md— update authoring and local run instructions for auto-registration. - [ ]
docs/frds/README.md— add FRD 0004 to the index. - [x] Evolution #112: update
docs/triggers.md,docs/workflows.md, anddocs/architecture.mdfor trigger-started workflows. - [x] Evolution #117: document
workflows.subagentsand thesub_agenttask indocs/front-matter-spec.md,docs/workflows.md, anddocs/architecture.md; keep the sample customer-facing and free of FRD/Durable implementation details. - [x] Evolution #151: document multi-agent policy isolation, workflow-ID
migration, final-agent drain operations, and the runnable
samples/per-agent-workflows/app. - [x] Evolution #151: update
samples/README.mdwith the multi-agent runnable sample while keeping internal verifier details outside the customer app. - [ ] Evolution #1276: update
docs/workflows.mdwithwhen,for_each, iteration locals, fan-in, limits, stable failures, and status examples. - [ ] Evolution #1276: update
docs/architecture.mdfor runtime materialization, deterministic scheduling, structured status hand-off, runtime-owned versus project-discovered skills, and direct/delegated capability paths. - [ ] Evolution #1276: update the selected workflow sample and its README with a collection-driven fan-out/fan-in scenario.
- [ ] Evolution #DTS display names: update
docs/workflows.mdanddocs/architecture.mdwith derived orchestration/Activity labels and the native orchestrator execution contract. - [ ] Evolution #1278 slice 1: update
docs/workflows.md,README.md, anddocs/architecture.mdfor plan-authored retry, the failure contract, idempotency, replay compatibility, and deferred decorator integration. - [ ] Evolution #1278 slice 1: add a runnable plan-authored retry sample and list
it in
samples/README.md. - [x] Timeout PR: document
execution.timeout,@workflow_tool(timeout=...), per-field precedence, the one-attempt default,workflow_task_timeout, and work that can continue after a deadline indocs/workflows.mdanddocs/architecture.md. Distinguish the library admission cap fromfunctionTimeoutand link to the hosting-plan limits. Explain the diagnostic warning, host logs, Application Insights, and configuration overrides. Include the timeout sample insamples/README.md. - [x] Continuation PR: document
execution.continue_on_error, permitted failure kinds, bounded failure results, and failure/cancellation order indocs/workflows.mdanddocs/architecture.md. Explain that completion does not imply task success and that the current UI does not distinguish continued failures. Include the continuation sample insamples/README.md.
8. Status & sign-off¶
- Architecture review (phase 2): Completed by
frd-reviewer(rubber-duck), 2026-07-06. Initial findings around pipeline boundaries, dual-decorator semantics, duplicate-name scope, non-main behavior, reserved names, and unknown decorator kwargs were addressed. Re-review found no remaining blocking issues and deemed the FRD ready for human sign-off. - Human sign-off: TsuyoshiUshio, 2026-07-06 →
status: Finalized. - Evolution review: Markdown-declared trigger support reviewed by TsuyoshiUshio and Chris Gillum in PR #112, 2026-07-23.
- Workflow Sub Agent architecture review: External contract reviewed in PR #117. Chris Gillum recommended direct Activity execution because current Serverless Agent invocations are stateless; the plan was revised to remove child orchestration and child ids. A dedicated pre-implementation review on 2026-07-24 additionally required an executable E2E sample, runtime authorization enforcement, explicit at-least-once semantics, and an Activity-owned timeout boundary; those findings are incorporated above.
- Workflow Sub Agent human sign-off: TsuyoshiUshio, 2026-07-24. Approved
Activity-only execution,
{agent, text}results, main-only v1 scope, and implementation using TDD followed by sample E2E validation. - Multi-agent workflow sign-off: TsuyoshiUshio, 2026-08-10. Approved app-wide Durable registration, immutable per-agent policies, agent/session isolation, deployed-policy Activity reauthorization, the experimental workflow-ID migration, and Storage/DTS E2E validation.
- Multi-agent architecture review: An independent rubber-duck review on
2026-08-10 evaluated the extension against
main,docs/architecture.md, FRDs 0004 and 0007, issues #1274/#1275, and the existing trigger and Workflow Sub Agent implementation. Findings on digest strength, provisional decisions, legacy-ID behavior, and verifier prerequisites were incorporated with no remaining blockers. The human sign-off ratified the resulting authorization, non-HTTP management, and digest decisions before implementation. - Final-agent drain sign-off: TsuyoshiUshio, 2026-08-12. Approved explicit runtime retention so pending work reaches Activity reauthorization instead of being stranded after the last workflow-enabled agent is removed.
- Dynamic control flow extension: Drafted for planning issue #1276 on 2026-08-13. An independent architecture review identified skip propagation, numeric instance ordering, and controlled-failure provenance as blocking ambiguities; this draft now defines each explicitly and also clarifies static serialization, iteration-local parsing, status schema versioning, and iterated waits. A final independent review found no blocking issues and deemed the extension ready for human review.
- Dynamic control flow human sign-off: TsuyoshiUshio, 2026-08-14. Approved
Decisions 43-53 as proposed and authorized implementation and testing. FRD
status returned to
Finalized. - Task execution policy decomposition: TsuyoshiUshio, 2026-08-29. Approved landing Durable native retry first, followed by per-attempt timeout with continue-on-error, then retry/timeout observability and structured status.
- Native retry architecture review: An independent rubber-duck review on 2026-08-29 confirmed the Durable native driver and raise-versus-return failure split. Its forward-compatibility and replay-safety findings are resolved by Decisions 75 and 76.
- Execution-foundation split approval: TsuyoshiUshio, 2026-09-02. Approved
extracting the complete plan-authored execution foundation from PR #185 while
leaving that PR and branch untouched; PR #185 will be rebased after this slice
merges to retain only
@workflow_tool(retry=...)integration. - Execution-foundation split review: An independent rubber-duck review on
2026-09-02 evaluated the split against the finalized FRD, architecture module
boundaries, replay compatibility, and PR #192 split rules. The review required
the delivery plan, corrected retry-window wording, explicit exclusion of
decorator integration, and a closed transient classification for Workflow Sub
Agent timeouts; Decisions 74, 78, and 79 and the design above resolve those
findings. FRD status remains
Finalized. - Execution-foundation final review: A read-only branch review on 2026-09-02
identified Durable's wall-clock evaluation of finite
retry_timeoutduring history replay. Decision 80 removes that nondeterministic input while retaining the submission-time retry-delay bound. FRD status remainsFinalized. - Decorator retry integration approval: TsuyoshiUshio, 2026-09-08. Requested the smallest current-main change equivalent to the rebased remainder of PR #185 after PR #193 merged.
- Decorator retry architecture review: An independent review on 2026-09-08 rejected rebasing the stale full branch because it would regress PR #193 contracts, and approved a residual-only port preserving persisted-input replay behavior, filtered immutable policy catalogs, and decorator-over-plan precedence. Decision 81 records the resulting scope.
- Timeout and continuation reimplementation approval: TsuyoshiUshio,
2026-09-10. Reviewed the stale PR #186 stack against current
main, confirmed the retry foundation had already landed through PRs #193 and #207, and directed a clean reimplementation of the remaining timeout and continue-on-error slice rather than a rebase. Decision 82 records the scope. - Timeout and continuation architecture review: An independent rubber-duck
review on 2026-09-10 evaluated the slice against current
main. It raised four blocking findings: collecting a whole wave would delay a non-continuable failure behind unrelated work; downgrade with new histories in flight was unaddressed; the continuable kinds and the continued-node result were underspecified, risking a quadratic payload from a reused failure envelope; and the test plan omitted mixed waves, cancellation races,for_eachfailure shapes, and real-host validation. All are resolved by Decisions 88 and 92-94 and the expanded §6. Non-blocking findings on Sub Agent timeout precedence, worker cancellation wording, the per-instance scope of the admission cap, field-wise decorator precedence, error-code mapping, and slice sizing are resolved by Decisions 85, 90, 95, and 96 and the design above. - Requirement review and human approval: TsuyoshiUshio, 2026-09-14.
Approved REDUCE with decorator timeout retained. Deliver timeout and
continuation as two stacked implementation PRs, each with tests and docs.
Use the existing transient failure kind with a separate timeout error code.
Preserve the old scheduler path for waves without enabled continuation and
define one attempt when retry is omitted. Decisions 97-100 record this
revision and replace the affected earlier decisions. The earlier review's
claim that all failures were already immediate was incorrect: returned
terminal outcomes are processed after the wave. The approved design now
distinguishes that path from Durable exceptions. Status remains
Finalized; implementation has not started in PR #212. - Final review delivery approval: TsuyoshiUshio, 2026-09-16. After both implementation slices and their tests were complete, directed one final PR for review. Decision 104 replaces only the two-PR review boundary from Decision 97. The timeout, continuation, and deferred status/UI scopes remain unchanged.
- Continuation testing review: An independent testing review on 2026-09-16 required explicit coverage for timeout continuation, opaque failures, dynamic retry exhaustion, aggregate status, wave timing, replay, and the cancellation race. The implementation and real-host E2E now cover these cases.