Skip to content

Commit 32c6676

Browse files
committed
Start only selected matrix interpreters and repair crashed runners
A plain `run` started every interpreter instance of a matrixed env while the run itself only executed the selected subset, so unselected children could fail to start (or be repaired) for no reason. Separately, an ER that died before publishing its port surfaced as a raw start error even though a reinstall usually fixes it. - Thread the run's env/interpreter selection into the payload-schema fetch (`runOptions`, honoured only with `startRunners`). The CLI builds one dict that feeds both the fetch and the run, so they cannot drift. Selection is computed only for projects with a matrixed action and selectors are never validated in the fetch, so a selector valid in a sibling project cannot fail it. - Repair a runner that crashed before its port (ServerExitedBeforePort in the __cause__ chain) the same way as NO_VENV: install, then restart. Only at the run gate and dispatch start, never during metadata resolution, and never for timeouts, which are load problems. - Rename repair_no_venv_env to repair_env and serialize repairs per (project, env) so two callers never install into one venv at once. - Pass the env's configured interpreter to the create_envs step in install_env_for_project so a repair rebuilds the venv with the right Python.
1 parent b594825 commit 32c6676

20 files changed

Lines changed: 1055 additions & 75 deletions

‎docs/cli.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -69,7 +69,7 @@ python -m finecode run [options] <action> [<action> ...] [payload] [--config.<ke
6969

7070
In a multi-project workspace, `run` fans out across every project that declares the action; spawned subprocesses are bounded by the machine-wide process budget (default: derived from the machine's CPU budget). Fan-out is throttled, never refused — workspace size does not limit which actions you can run. See [Process budget](guides/wm-server-internals.md#process-budget).
7171

72-
`--env` and `--interpreter` on `run` use the same selector semantics as `prepare-envs` (ADR-0050): they compose by intersection, and a matrix env's config-declared `default_interpreters` policy (see [Preparing Environments — default interpreter subset](guides/preparing-environments.md#default-interpreter-subset)) applies as the default when neither is given — so a plain `run` can execute only a local subset of a matrix (e.g. the newest interpreter) while CI still runs the full axis, mirroring `prepare-envs`.
72+
`--env` and `--interpreter` on `run` use the same selector semantics as `prepare-envs` (ADR-0050): they compose by intersection, and a matrix env's config-declared `default_interpreters` policy (see [Preparing Environments — default interpreter subset](guides/preparing-environments.md#default-interpreter-subset)) applies as the default when neither is given — so a plain `run` can execute only a local subset of a matrix (e.g. the newest interpreter) while CI still runs the full axis, mirroring `prepare-envs`. The selection also decides which interpreter instances are *started*, not only which run: unselected matrix children are never started or repaired.
7373

7474
WAL environment variable and storage settings are shared with `start-wm-server` — see [`start-wm-server`](#start-wm-server) for details.
7575

‎docs/guides/preparing-environments.md‎

Lines changed: 9 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -224,6 +224,15 @@ The ER signals the problem by returning error code `-32001` (`ENV_REINSTALL_NEED
224224

225225
The WM catches this, runs `CreateEnvsAction` + `InstallEnvsAction` for the affected env, then restarts the ER.
226226

227+
### Other triggers
228+
229+
The same install-then-restart repair also runs when a runner the run needs fails to start:
230+
231+
- `NO_VENV`: the venv is missing (or was just wiped as stale/relocated). Repaired wherever the failure surfaces — the run gate, the dispatch start, and metadata resolution.
232+
- Crash before the port: the ER process exited before publishing its port (`ServerExitedBeforePort` in the failure's `__cause__` chain). Repaired only for an env the run needs — the run gate and the dispatch start — never during metadata resolution, and never for unselected matrix children, which the gate does not start.
233+
234+
Timeouts (the port wait expiring while the process is still alive) are load problems, not broken venvs, and are never repaired. Every repair runs at most once per start attempt: if the restart still fails, the error names the env and the project. Concurrent repairs of the same env are serialized so two callers never install into one venv at the same time.
235+
227236
### Runner routing
228237

229238
The runner that executes `CreateEnvsAction` / `InstallEnvsAction` during auto-repair depends on which env is being fixed:

‎docs/guides/wm-server-internals.md‎

Lines changed: 4 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -545,7 +545,10 @@ Triggered when `runner_manager.update_runner_config` receives error `-32001` fro
545545
sequence as `prepare-envs` — scoped to the affected env — then restarts the ER via
546546
`restart_extension_runner`. The only difference from a manual run is which runner executes
547547
those actions; see [Automatic env repair](../guides/preparing-environments.md#automatic-env-repair)
548-
for the routing rules.
548+
for the routing rules. The same repair also covers two startup failures of an env
549+
the run needs: a missing venv (`NO_VENV`) and an ER that crashed before publishing
550+
its port. Timeouts and unselected matrix children are never repaired, and each
551+
repair runs at most once per start attempt.
549552

550553
---
551554

‎docs/wm-protocol.md‎

Lines changed: 9 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -432,6 +432,15 @@ probes them again. The CLI sets it so payload values with a path type can be
432432
converted before dispatch; MCP leaves it off so listing tools never starts
433433
environments.
434434

435+
`runOptions` (optional, honoured only with `startRunners`): the selection
436+
inputs the run itself will use — `devEnv`, `envSelectors` and
437+
`interpreterSelectors` (same shapes as the `actions/runBatch` run options).
438+
The fetch computes the same per-project interpreter selection as the run, so
439+
it starts only the interpreter instances the run will select. Selectors are
440+
never validated here: a selector valid in another in-scope project but not in
441+
the schema project must not fail the fetch; the run's own validation reports
442+
bad selectors a moment later.
443+
435444
**Result:**
436445

437446
```json

‎finecode_jsonrpc/src/finecode_jsonrpc/__init__.py‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -5,6 +5,7 @@
55
NoResponse,
66
RequestCancelledError,
77
ResponseTimeout,
8+
ServerExitedBeforePort,
89
ServerFailedToStart,
910
ServerStoppedError,
1011
StartupTimeline,
@@ -30,6 +31,7 @@
3031
"NoResponse",
3132
"RequestCancelledError",
3233
"ResponseTimeout",
34+
"ServerExitedBeforePort",
3335
"ServerFailedToStart",
3436
"ServerStdioTransport",
3537
"ServerStoppedError",

‎src/finecode/cli_app/commands/run_cmd.py‎

Lines changed: 11 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -377,6 +377,13 @@ async def _attach_session(*, first_connect: bool) -> None:
377377
schema_project = _choose_schema_project(
378378
project_paths, all_actions, action_sources, workdir_path
379379
)
380+
# The schema fetch and the run must select the same interpreter
381+
# instances: one dict feeds both, so they cannot drift.
382+
selection_options = {
383+
"devEnv": dev_env,
384+
"envSelectors": env_selectors or [],
385+
"interpreterSelectors": interpreter_selectors or [],
386+
}
380387
action_payload = await _resolve_payload(
381388
client=client,
382389
action_payload=action_payload,
@@ -385,6 +392,7 @@ async def _attach_session(*, first_connect: bool) -> None:
385392
schema_project=schema_project,
386393
base_dir=workdir_path,
387394
map_payload_fields=map_payload_fields,
395+
run_options=selection_options,
388396
)
389397

390398
# Workspace-scoped actions run once on the root project and stream all
@@ -420,16 +428,14 @@ async def _attach_session(*, first_connect: bool) -> None:
420428
"concurrently": concurrently,
421429
"resultFormats": result_formats,
422430
"trigger": "user",
423-
"devEnv": dev_env,
431+
**selection_options,
424432
# Ask the WM to type-safely merge streamed partials per project/action
425433
# and return the merged result, so the returned/saved data is complete
426434
# even when one project streams many partials.
427435
"mergeResults": True,
428436
# PRD-0003 AC8: WM-only selectors restricting a matrixed
429437
# action's fan-out to a subset of its declared interpreter axis.
430438
# Never forwarded to an ER.
431-
"envSelectors": env_selectors or [],
432-
"interpreterSelectors": interpreter_selectors or [],
433439
}
434440

435441
partial_result_token = str(uuid.uuid4())
@@ -619,6 +625,7 @@ async def _resolve_payload(
619625
schema_project: str,
620626
base_dir: pathlib.Path,
621627
map_payload_fields: set[str] | None,
628+
run_options: dict[str, typing.Any] | None = None,
622629
) -> dict[str, typing.Any]:
623630
"""Build the final payload from the raw CLI strings and the payload schemas.
624631
@@ -643,7 +650,7 @@ async def _resolve_payload(
643650
# action, so a path value can be converted before dispatch.
644651
try:
645652
schemas = await client.get_payload_schemas(
646-
schema_project, action_sources, start_runners=True
653+
schema_project, action_sources, start_runners=True, run_options=run_options
647654
)
648655
except ApiError as exc:
649656
raise RunFailed(

‎src/finecode/wm_client.py‎

Lines changed: 9 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -16,6 +16,7 @@
1616
import json
1717
import pathlib
1818
import random
19+
import typing
1920

2021
from loguru import logger
2122

@@ -447,6 +448,7 @@ async def get_payload_schemas(
447448
action_sources: list[str],
448449
*,
449450
start_runners: bool = False,
451+
run_options: dict[str, typing.Any] | None = None,
450452
) -> dict[str, schema_utils.PayloadSchema | None]:
451453
"""Return payload schemas for the given actions in a project.
452454
@@ -459,6 +461,11 @@ async def get_payload_schemas(
459461
environments before probing so a schema that needs one is
460462
available. Defaults to false so passive listing never starts
461463
environments.
464+
run_options: Selection inputs honoured only with
465+
``start_runners``: ``devEnv``, ``envSelectors`` and
466+
``interpreterSelectors`` — the same values the run itself
467+
will use, so the fetch starts only the interpreter
468+
instances the run will select.
462469
463470
Returns:
464471
Mapping of action source → JSON Schema fragment, or ``None``
@@ -467,6 +474,8 @@ async def get_payload_schemas(
467474
params: dict = {"project": project, "actionSources": action_sources}
468475
if start_runners:
469476
params["startRunners"] = True
477+
if run_options is not None:
478+
params["runOptions"] = run_options
470479
result = await self.request(
471480
"actions/getPayloadSchemas",
472481
params,

‎src/finecode/wm_server/_api_handlers/_actions.py‎

Lines changed: 42 additions & 14 deletions
Original file line numberDiff line numberDiff line change
@@ -39,19 +39,11 @@ async def _handle_run_action(
3939
parsed = await _parse_and_validate_run_action_params(raw_params, ws_context)
4040

4141
try:
42-
# Start required environments so that canonical_source is populated for all
43-
# importable action classes (ADR-0021: canonical_source is the WM's internal
44-
# identifier after runner initialization; only find_action_by_source at the
45-
# external API boundary may match by config source alias).
46-
await run_service.start_required_environments(
47-
{parsed.project.dir_path: [parsed.action.name]},
48-
ws_context,
49-
initialize_all_handlers=True,
50-
)
5142
# PRD-0003 AC8: resolve `--env`/`--interpreter` selectors
5243
# (+ config default) for this project, to restrict a matrixed
5344
# action's fan-out. `None` (no selectors, no narrowing default)
54-
# runs the full declared axis, unchanged.
45+
# runs the full declared axis, unchanged. Resolved before the
46+
# gate so the gate starts only the children the dispatch runs.
5547
run_selection.validate_run_selectors(
5648
parsed.options.get("envSelectors", []),
5749
parsed.options.get("interpreterSelectors", []),
@@ -65,6 +57,18 @@ async def _handle_run_action(
6557
parsed.dev_env.value,
6658
ws_context,
6759
)
60+
# Start required environments so that canonical_source is populated for all
61+
# importable action classes (ADR-0021: canonical_source is the WM's internal
62+
# identifier after runner initialization; only find_action_by_source at the
63+
# external API boundary may match by config source alias).
64+
await run_service.start_required_environments(
65+
{parsed.project.dir_path: [parsed.action.name]},
66+
ws_context,
67+
initialize_all_handlers=True,
68+
selected_interpreters_by_project={
69+
parsed.project.dir_path: selected_interpreters
70+
},
71+
)
6872
executor = run_service.ProjectExecutor(ws_context)
6973
result = await executor.run_action(
7074
action_source=parsed.action.canonical_source,
@@ -428,12 +432,36 @@ async def _probe_handler_envs(names: list[str]) -> None:
428432
# startup tool listing never starts every handler env.
429433
start_runners = params.get("startRunners", False)
430434
if start_runners and still_missing:
435+
from finecode.wm_server.config import env_selection
436+
from finecode.wm_server.config.interpreter_matrix import (
437+
InvalidInterpreterError,
438+
)
431439
from finecode.wm_server.services import run_service
440+
from finecode.wm_server.services.run_service import run_selection
432441

433-
await run_service.start_required_environments(
434-
{project.dir_path: still_missing}, ws_context
435-
)
436-
await _probe_handler_envs(still_missing)
442+
run_options = params.get("runOptions") or {}
443+
try:
444+
selection = run_selection.selection_for_matrixed_actions(
445+
{project.dir_path: still_missing},
446+
run_options.get("envSelectors", []),
447+
run_options.get("interpreterSelectors", []),
448+
run_options.get("devEnv", "cli"),
449+
ws_context,
450+
)
451+
except (env_selection.EnvSelectionError, InvalidInterpreterError) as exc:
452+
# Never validate selectors here: the run validates against
453+
# every in-scope project and tolerates a selector valid in
454+
# only one of them, while this fetch sees a single project.
455+
# The run's own validation raises the real error a moment
456+
# later; the schemas stay None until then.
457+
logger.debug(f"Skipping runner start for schema fetch: {exc}")
458+
else:
459+
await run_service.start_required_environments(
460+
{project.dir_path: still_missing},
461+
ws_context,
462+
selected_interpreters_by_project=selection,
463+
)
464+
await _probe_handler_envs(still_missing)
437465

438466
# Re-key schemas by the requested source rather than internal action name.
439467
result_schemas: dict[str, dict | None] = {}

‎src/finecode/wm_server/_api_handlers/_streaming.py‎

Lines changed: 17 additions & 13 deletions
Original file line numberDiff line numberDiff line change
@@ -341,7 +341,22 @@ async def _handle_run_batch_with_partial_results(
341341
f"{sorted({a for names in actions_by_project.values() for a in names})}"
342342
)
343343

344-
await run_service.start_required_environments(actions_by_project, ws_context)
344+
try:
345+
selection_by_project = run_selection.selection_for_matrixed_actions(
346+
actions_by_project,
347+
parsed.env_selectors,
348+
parsed.interpreter_selectors,
349+
parsed.dev_env.value,
350+
ws_context,
351+
)
352+
except env_selection.EnvSelectionError as exc:
353+
raise ActionRunFailed(str(exc)) from exc
354+
355+
await run_service.start_required_environments(
356+
actions_by_project,
357+
ws_context,
358+
selected_interpreters_by_project=selection_by_project,
359+
)
345360

346361
payload_overrides = parsed.params_by_project or {}
347362
# Lock to prevent concurrent writes to the shared writer from project tasks.
@@ -377,18 +392,7 @@ async def _stream_action(
377392
else None
378393
)
379394
if action_def is not None and matrix_runner.is_matrixed(action_def):
380-
try:
381-
selected_interpreters = (
382-
run_selection.selected_interpreters_for_project(
383-
project_path,
384-
parsed.env_selectors,
385-
parsed.interpreter_selectors,
386-
parsed.dev_env.value,
387-
ws_context,
388-
)
389-
)
390-
except env_selection.EnvSelectionError as exc:
391-
raise ActionRunFailed(str(exc)) from exc
395+
selected_interpreters = selection_by_project.get(project_path)
392396

393397
async def _on_partial(
394398
interpreter_canonical: str, result_by_format: dict

‎src/finecode/wm_server/context.py‎

Lines changed: 9 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -210,6 +210,15 @@ class WorkspaceContext:
210210
# from inside that same Phase 2.
211211
env_install_locks: dict[Path, asyncio.Lock] = field(default_factory=dict)
212212

213+
# Per-env repair locks. Guard per-env repair (install + restart) so two
214+
# concurrent repairs of the same (project, env) never install into one
215+
# venv at the same time. Keyed per env and separate from
216+
# env_install_locks: install_env_for_project takes env_install_locks[project]
217+
# inside the repair, so sharing that dict would self-deadlock.
218+
env_repair_locks: dict[tuple[Path, str], asyncio.Lock] = field(
219+
default_factory=dict
220+
)
221+
213222
# Both budgets below are sized from ONE combined machine budget in
214223
# __post_init__ (ADR-0093): their sum stays at or below
215224
# machine_subprocess_budget(), so the WM's event loop keeps a free core.

0 commit comments

Comments
 (0)