Skip to content

Improve concurrency handling, fix part of CI pipeline and more - #39

Open
Aksem wants to merge 116 commits into
mainfrom
fix/lsp-request-concurrency-cap
Open

Aksem wants to merge 116 commits into
mainfrom
fix/lsp-request-concurrency-cap

Conversation

@Aksem

@Aksem Aksem commented Sep 5, 2026

Copy link
Copy Markdown
Member

No description provided.

Aksem added 30 commits July 21, 2026 06:56
…tartup fan-out (ADR-0063)

Three related gaps in how the WM manages Extension Runner processes:

1. A start attempt that spawns the OS process but then fails or times
   out before the port handshake/RPC connection completes (debugger
   port wait, connect_to_server) left that process orphaned — nothing
   ever revisited a runner once it was marked FAILED. JsonRpcClient
   gains force_kill() (SIGKILL via os.killpg, since the process is
   started with start_new_session=True so it and anything it spawned,
   e.g. a package-manager invocation, can be reached as one group);
   every failed-start path in _start_extension_runner_process now
   calls it. client.pid is now captured as soon as the process is
   spawned (before the handshake), so force_kill has a target even
   when start() never returns successfully.

2. shutdown_service.on_shutdown stopped RUNNING/REPAIRING runners one
   at a time via the old *_sync helpers, so total shutdown time scaled
   with runner coun
   bounded to _STOP_TIMEOUT_SEC, but the total no longer multiplies.
   A runner still INITIALIZING at shutdown was never sent a graceful
   exit (that only happens once RUNNING) and has no way to learn the
   WM is going away, so those are force-killed directly instead.
   The now-unused stop_extension_runner_sync/shutdown_sync/exit_sync
   helpers are removed. Deliberately no force-kill fallback was added
   to stop_extension_runner's own timeout path: an ER that already
   received exit may still be tearing down its own spawned
   subprocesses, and killing it mid-cleanup risks orphaning exactly
   the children a slower-but-graceful exit would have reaped itself.

3. Nothing bounded how many ERs could be starting at once — process
   spawn + interpreter init + imports is CPU/memory-bursty, and every
   trigger (workspace init, a matrixed run, prepare-envs' runner-start
   step) funnels through the same _start_extension_runner_process
   chokepoint. It's now held under ws_context.er_startup_semaphore
   (sized via FINECODE_WM_MAX_CONCURRENT_ER_STARTS or the shared
   machine-budget formula, ADR-0063) from just before spawn until the
   RPC channel connects — not around whatever the triggering action
   does afterward in the ER's own process, a separate and far more
   variable cost this cap deliberately leaves unconstrained.

Also switches the `@typing.override` decorator to a version-gated
`typing`/`typing_extensions` import in the two lowest-level protocol
files touched here (finecode_jsonrpc/client.py,
wm_server/runner/_internal_client_types.py), since `typing.override`
requires 3.12+ and these run under the workspace's >=3.11 floor;
pulls in typing-extensions as an explicit dependency where needed.
…ncing (PRD-0004-AC6/AC7)

An otlp_endpoint that's merely malformed previously surfaced as a
confusing failure deep inside OTel SDK setup. telemetry._validate_endpoint
now parses host+port up front (with or without a scheme) and raises a
clear, actionable error at the three init_* call sites (logging, tracer,
meter provider) before touching the SDK.

A configured-but-unreachable collector — the common case when FineCode
starts before `scripts/observability.sh up`, or before a Compose otel
profile comes up alongside the container — no longer needs to be reachable
at startup: exporters already buffer and retry, so this doesn't change
behavior, but a stream of gRPC export-retry warnings for the process
lifetime is now silenced (opentelemetry.exporter logger raised to ERROR)
in favor of a single one-time reachability heads-up per endpoint, logged
once and cached in _probed_endpoints so it can't repeat.

Documents the resulting contract in configuration.md: malformed values
fail fast, unreachable ones don't block startup, and points to the
Observability guide for running a backend locally.
Ships the core release-automation loop: release_workspace_packages
(fine_release) discovers every workspace package whose declared version
is absent from its registry, computes a dependency-respecting publish
order via fine_dep_graph, and sweeps them in order. A package whose
transitive same-run dependency failed is BLOCKED rather than attempted
(AC7, AC11); a successfully published package gets a best-effort git
tag recording the publish (ADR-0060), via a new minimal fine_git preset
(create_git_tag, push_git_refs) that fine_release depends on without
pulling in release semantics.

Getting there required fixing a registry-endpoint bug the old flow
never exercised: Repository carried a single `url`, but PyPI splits
reads and writes across two hosts (pypi.org vs upload.pypi.org). Reusing
one URL for both silently misroutes — an index lookup against the
upload host 404s, which is indistinguishable from "nothing published"
and would make a package look unpublished when it isn't. Repository
now carries index_url/upload_url explicitly (matching pip's/twine's own
naming), and fine_python_package_info.registry_endpoints validates each
is used for its own role before any request goes out.

is_artifact_published_to_registry is replaced by list_published_artifacts,
which reports the registry's actual filenames for a version instead of a
bool per caller-supplied dist path. This lets a caller (the release
sweep, a dry-run preview) ask "what's published" before it has built
anything locally, and derive what still needs uploading by filename
membership rather than needing to already know the dist paths up front.

publish_artifact_to_registry and publish_artifact now report registry
failures in their results (`error`/`failed_registries`) instead of
raising, and publish_artifact_handler dispatches to registries with
asyncio.gather instead of a TaskGroup — one registry's failure no
longer cancels uploads already in flight to the others (ADR-0062).
publish_and_verify_artifact's result gains a matching `publish_errors`
map alongside its existing `verification_errors`.

get_src_artifact_version's src_artifact_def_path is now optional,
defaulting to the current project, so release handlers fanning the
same payload out to every candidate project don't each need to resolve
their own path first.

Adds finecode_extension_runner/testing: an in-memory handler test
harness (run_handler/handler_test_session) that boots a real ER
in-process against stub services (IFileEditor, ICommandRunner, WAL
writer, progress/partial-result senders) instead of stubbing each
handler's dependencies by hand per test. Used throughout the new
fine_git and fine_python_package_info tests, and covers its own
progress-collection behavior.

CI workflow env vars are updated to the new index_url/upload_url shape
for both TestPyPI and PyPI publish steps; also drops a stale
"test with all supported python versions" TODO that predates the
interpreter-matrix work already handling that.
… (ADR-0067)

Two independent gaps closed together:

Converting an env to an interpreter matrix (ADR-0047) silently orphans
its old venv: config only sees the expanded children afterward, so
neither prepare-envs nor --recreate ever revisits the base name's
.venvs/ directory again. Same fate for a renamed env or a preset that
stops contributing one — the venv just sits there, easily hundreds of
MB per project. list_envs reports declared/on-disk/orphaned status per
env from the filesystem alone (cheap across a whole workspace, no
interpreter execution); remove_envs deletes them, defaulting to the
orphaned set. Two guards apply to whatever was asked for before
existence is even checked, so a rejection can't depend on disk state:
a declared env needs force (an ER may be running in it), and the
current env is never removable regardless. IFileManager.remove_dir
gains tolerant=True for this: read-only contents, dangling symlinks,
and already-absent paths are all treated as "make it gone" rather than
errors, since a half-created or permission-damaged venv is exactly
what removal should clear.

Separately, `run` fan-out across projects had no concurrency bound,
unlike prepare-envs (ADR-0055) — a workspace-wide run put every
project's ER to work at once, each free to spawn subprocesses up to
its own ICommandRunner cap, so the two layers compose multiplicatively
exactly as ADR-0055 warned about. max_project_fanout couldn't be
reused directly to fix this: it's a runaway-recursion guard that
refuses, meant for nested orchestration (ADR-0016), and applying it at
depth 0 would make every workspace-wide action unusable past an
arbitrary workspace size someone chose on purpose. Depth 0 now gets a
per-call semaphore instead (FINECODE_WM_RUN_MAX_CONCURRENT_PROJECTS,
same sqrt-split default formula as ADR-0055) that throttles without
ever refusing; the recursion guard still applies, and only applies,
once orchestration_depth > 0. The semaphore is built fresh per fan-out
call rather than shared, since a shared one would let an outer fan-out
hold every permit while waiting on an inner one that can never
acquire any.
fine_release's release_workspace_packages action previously owned both
the cross-package concerns (candidate discovery, dependency order,
blocking, tagging, pushing) and the single-package concerns (build,
publish, verify) in one handler chain, so a package could not
customize its own release steps without touching the workspace-wide
action. ReleasePackageAction is now a package's own chain
(build/publish/record_tag), and release_workspace_packages delegates
to it once per candidate via run_action_in_projects, keeping only what
is inherently cross-package: ordering, dependent-blocking, and
publishing every git ref the run produced in a single push (fine_git's
push_git_refs) rather than once per package.

Extends ADR-0054's canonical-source resolution (previously actions
only) to handlers: resolve_action_meta now reports canonicalSource per
handler alongside fileLoc, since config almost always declares a
handler via its package's __init__.py re-export rather than the
module the class is actually defined in. ActionHandler gains a
canonical_source field (None until the hosting ER resolves it, or
permanently if the class can't be imported), and
parse_workspace_actions carries it through from the wire payload.

finecode_extension_runner/testing's handler_test_session needed a way
to observe progress calls directly instead of only through the
token/global-function forwarding meant for production RPC, so
run_action now accepts an explicit progress_sender that bypasses that
path. Also fixes a bug the new fine_release tests exposed: handler
config structuring only caught cattrs.ClassValidationError, so a
malformed entry in a list- or dict-typed config field raised the
sibling IterableValidationError uncaught instead of surfacing as a
readable ActionFailedException.

create_envs now reports whether it actually built a virtualenv or
found a valid one already there (CreateEnvsRunResult.created), and
create/install env progress messages use a new env_label() helper
(<project>/<env>) instead of the bare env name, which was ambiguous
whenever one dispatch call spans multiple same-named envs across
projects (e.g. prepare-envs' dev_workspace bootstrap step).
remove_dir's --recreate path now passes tolerant=True, since a
half-created or permission-damaged venv being removed before
recreation is exactly the case tolerant removal exists for.

Clarifies IProjectInfoProvider.get_project_raw_config's docstring: the
config it serves is resolved (presets merged, interpreter matrices
already expanded per ADR-0047), not the file's own raw content, which
was easy to misread as unprocessed. Adds a regression test pinning
that the expansion happens before the config is stored for serving,
not just that resolve_interpreter_matrices itself produces the right
shape, plus a comment in prepare_envs_service.py on why step 3 keeps
installing into dev_workspace even though step 5 excludes it.
…070)

A service declaration's identity is its `interface` dotted path, but that
can't be written readably in an environment variable (dots collapse to
underscores and collide with the `__` nesting separator already used for
handler config). FINECODE_SERVICE_CONFIG_<NAME>__<PARAM_PATH> now
addresses a service by a short alias instead: derived from the
interface's class name (leading `I` + uppercase stripped, snake-cased),
never declared, so a declaration and an environment can never disagree
about it. ServiceConfigResolver (finecode_extension_runner/service_config.py)
merges raw_config (activator) < declared config < env override by deep
merge, and reports ambiguous aliases (two interfaces deriving the same
name) or unmatched override names once all bindings have registered.
Resolution happens in the Extension Runner, not the WM, because only the
ER can see activator-registered bindings alongside declared ones.

IServiceRegistry.register_impl drops its `singleton` flag: every binding
is already cached for the runner's life, and the flag never controlled
that -- it only masked the case where a `[[tool.finecode.service]]` entry
had no way to pass it, silently handing a concrete-injecting handler a
second instance. register_impl now always alias-binds the concrete type
to the interface's instance.

Config-only `[[tool.finecode.service]]` entries (no `source`) are now
legal: they carry config for a binding an activator already owns instead
of restating and pinning its implementation. IRepositoryCredentialsProvider
loses its push-seed methods (set_credentials/add_repository) since they
only existed to fake provisioning that belongs to the concrete impl
(ADR-0068); ConfigRepositoryCredentialsProvider is provisioned entirely
through service config now, and the old
PublishAndVerifyArtifactInitRepositoryProviderHandler -- an init action
that only carried config -- is deleted per S-203. The one remaining
init_repository_provider_handler is documented as dynamic-runtime-seeding
only, injecting the concrete provider type rather than the interface.

Adds docs/guides/designing-services.md: the S-1xx (contract) / S-2xx
(provisioning) / S-3xx (registration/lifecycle) rule set this change
implements, distilled from ADR-0038/0056/0068/0070. docs/configuration.md
documents the env-var format and service-name derivation;
docs/reference/services.md, docs/wm-er-protocol.md and
docs/guides/designing-actions-rules.md are updated to match.

HandlerInfo gains a `canonical_source` field (module a handler class is
actually defined in, vs. the config-facing `source` alias), following the
same canonical-source resolution already used for actions.

Also:
- Registers IKnowledgeStore (finecode_extension_api) with an ER-side
  KnowledgeStoreImpl, and wires runner_manager/wm_server to route two
  knowledge methods to it.
- Adds typing-extensions as an explicit dependency for fine_python_pyrefly,
  fine_python_ruff, fine_toml_tombi and finecode_extension_runner, since
  their `override` compat shim for Python <3.12 depends on it directly
  rather than transitively.
- Adds the fine_python_envs preset (derives an env's interpreter axis from
  requires-python via fine_envs' sync_toolchains/check_toolchains
  contracts) and a fine_git README.
- RuffFormatFileHandler now returns an empty `code` when the file did not
  change, instead of the unchanged content.
check_toolchains (ADR-0053) only surfaced axis drift through its own
run result; nothing routed it into audit_code, the umbrella CI/editor
diagnostics run through. CheckToolchainsAuditCodeBridgeHandler adds
that path: one ERROR diagnostic per stale env, anchored at the
project's definition file (position 0,0, matching how import-linter
anchors whole-project violations) since the drift belongs to the file,
not a line in it. CheckToolchainsRunResult gains project_def_path so
the bridge has something to key diagnostics on; CheckToolchainsHandler
fills it from IProjectInfoProvider since it runs in the project's own
ER and is the only place that knows the path without a second lookup.
The dependency on fine_audit_code/fine_inspect_code stays optional
(fine_envs[audit]) since fine_envs is the mandatory base preset every
project loads and only this one module needs the audit umbrella.

While building the bridge's tests, found the same bug already live in
check_toolchains_precommit_bridge_handler: it merged every project's
CheckToolchainsRunResult with update(), but that merge is scoped
within one project (R-302) and EnvToolchainAxis is keyed by env name,
unique inside a project but not across them. Two projects sharing the
`testing` env name (every matrix project does) collapsed into one
entry showing only the first project's versions, silently hiding the
second project's drift. Cross-project aggregation now happens in the
bridge itself, keeping one action_results entry per drifted project
under its own label.

Also:
- fine_python_test's run_tests moves from the single-interpreter `dev`
  env to a new `testing` matrix env whose interpreter axis is derived
  from requires-python via fine_python_envs' sync_python_interpreters
  (ADR-0053/0057), instead of running tests against only `dev`'s
  interpreter.
- Adds regression tests for _collect_services_in_config covering
  config-only entries (source/env omitted) and same-process alias
  collisions, which ADR-0070 left unmatched by test coverage.
read_file previously took a block=True flag shared with the same
_blocked_files map that a modifying caller used, so a read nested
anywhere inside an in-flight modification (e.g. formatting
pyproject.toml triggers a language-detection lookup that reads
pyproject.toml through a second session) waited on a release only its
own caller could perform, deadlocking the ER. Reads are now always
non-blocking and modification is its own operation, modify_file,
which claims exclusion on the file path (not the session) for as long
as the claim is held. Session identity is kept only for teardown:
ending a session releases whatever claims it still holds.

save_file gains an optional if_version to make a write conditional on
the version it was based on: if the file changed underneath the
caller, FileVersionConflict is raised and nothing is written, rather
than silently discarding whatever produced the newer version.
SaveFormatFileHandler uses this so a formatting pipeline whose input
went stale mid-run refuses the write and logs, instead of last-writer-
wins clobbering a concurrent edit.

InMemoryFileEditor's test double derives versions from content hash
instead of a constant "1", so if_version checks in tests mean
something, and gains matching modify_file locking.
LSP leaves parallel request execution to the server's discretion, so a client
that pipelines without bound relies on unspecified behaviour. tombi guards its
document, reference and schema stores with locks it reports contention on, and
with many documents in flight on one session individual requests stall past any
usable timeout while the process sits idle rather than busy.

Add an opt-in max_concurrent_requests to LspService, defaulting to unbounded so
pyrefly and ruff-lsp are unaffected, and set it to 1 for tombi. The limit bounds
whole interactions rather than single messages: the document synchronization
around a request mutates the same server state the request reads, so letting
another document's notifications interleave defeats the point. Forwarded file
events take the limit too, before the per-uri lock — the reverse order
deadlocks against format_file, which already holds them that way round.

Measured over 194 workspace TOML files against one tombi process at a 30s
timeout: unbounded fails 88 of 194 requests; every bounded value from 1 to 8
completes cleanly in 41-51s.

Alongside it:

- Drop didOpen for a change to a document the server does not hold open.
  didClose is only sent for documents an editor session opened, so nothing
  would ever close it again and every file written through the editor
  accumulated in the server for the lifetime of the session.
- Send didClose even when a format request fails, for the same reason.
- Answer workspace/inlayHint/refresh, which servers send regardless of the
  declared client capability — for some, once per document sync.
- Run tombi with --offline. No action here consumes live dependency data, so
  its per-dependency network lookups are latency for results nothing reads.

Claude-Session: https://claude.ai/code/session_019ZMjy2zw1jTjkfXCWNfEe6
Whitespace/line-wrap/quote-style normalization from running the
project formatter on every file with no pending unrelated changes.
Files that already had unstaged edits before formatting are excluded
from this commit so their formatting and content changes aren't
mixed together.
format and check_formatting disagreed persistently: format reported a
set of files as reformatted, check_formatting run straight afterwards
reported the same files as needing formatting, and neither run changed
what was on disk. Since check_formatting is format with save=False,
that can only mean the pipeline was not deterministic across runs.

IsortFormatFileHandler built isort.settings.Config without anchoring it
to the project, so isort inferred which packages are first-party from
the cwd of the process-executor worker -- a value the handler does not
control. With the project directory current, a project's own package is
first-party and its imports get their own trailing section; with any
other directory current, nothing is first-party and the same imports
collapse into the third-party block. Consecutive runs therefore
produced two different layouts for the same file and each run reported
the other one's output as needing formatting.

settings_path alone does not fix this: isort derives src_paths, which
is what first-party placement actually consults, from `directory`,
which independently defaults to os.getcwd(). Both are now set to the
project directory, so the same file sorts identically regardless of
where the runner was started.

The cwd was unstable because JsonRpcClient.start established it by
mutating the workspace manager's own cwd around the spawn and restoring
it afterwards. Runners start up to er_startup_semaphore at a time (7 on
this machine) and the spawn itself is dispatched to the io thread, so
an ER could be forked while another client held the cwd -- inheriting
that project's directory, or the restored original. The VIRTUAL_ENV
pop/restore had the same shape. Both are now passed to
create_subprocess_shell explicitly, which is what StdioTransport
already did.

Verified with a whole-workspace format followed immediately by
check_formatting: 43 files reformatted, then 0 files needing
formatting.

The loguru import move in client.py is that fix applied to this file --
it was one of the files whose import layout had been flipping.
Second pass of formatter normalization (import ordering, line
wrapping) over files not covered by the previous formatting commit.
Files with pre-existing unstaged edits are still excluded so their
formatting and content changes aren't mixed together.
The wm-layered import-linter contract was recently repaired to point at
real modules (it previously referenced ones that didn't exist, so it
never ran). Running it for the first time surfaced 7 direction
violations against the documented layer stack (docs/guides/wm-server-
internals.md). All are implementation debt, not a contract bug — fixed
in four groups:

- domain.py imported ErLoggingConfig from config.config_models (a
  stranded pure-data type). Moved it into domain.py; config_models.py
  and runner_client.py now import it from there.

- context.py importing runner.runner_client for ExtensionRunnerInfo
  can't be fixed the same way — that type wraps a live JsonRpcClient,
  so it can't move down to domain. Documented as an intentional
  exception in the wm-layered contract, mirroring the existing
  wm-domain-purity carve-out for context.py.

- Five call sites (_api_handlers/_workspace.py, services/
  runner_start_service.py, runner/runner_manager.py x3) reached back
  into wm_server.py/services via deferred imports inside function
  bodies to broadcast client notifications, forward ER logs, and
  dispatch ER-initiated action runs. Replaced with two new bridge
  modules (runner/wm_bridge.py, runner/run_dispatch_bridge.py),
  mirroring the existing runner/knowledge_bridge.py pattern: a
  Protocol + install/handlers slot living in the lower layer, filled
  by the higher layer at import time. The ER-dispatch handler bodies
  moved out of runner_manager.py into services/run_service/
  er_dispatch.py.

- config/read_configs.py called runner.runner_client directly to
  resolve py-preset install paths via an already-running dev_workspace
  runner. Split read_project_config into a pure read/merge pair
  (read_project_config_sources / finish_project_config) and moved the
  RPC-dependent preset resolution into a new runner/preset_resolution.py,
  which runner_manager.py (the only caller that ever needed it) now
  calls directly. No opaque object/Any typing needed anywhere in this
  path.

Verified with lint-imports (5/5 contracts kept, was 1 broken) and the
unit suite (87 failed/239 passed before and after — pre-existing
pytest-asyncio gap in this venv, confirmed via git stash comparison).
Ruff, black and isort each guessed a language level on their own --
ruff defaulted to py38, black had none configured at all -- so a
project's declared requires-python, its lint target, and its format
target could all name different Python versions with nothing to
notice the disagreement.

Add get_src_artifact_toolchain_range as the one place that range is
read (from requires-python for Python, via a new handler in
fine_python_package_info), shared with the interpreter axis derivation
in sync_python_interpreters so both read requires-python the same way.
Black and ruff now derive their target version from it when not
explicitly configured; an explicit value still wins.

Ruff's target version and format/lint settings are pushed to its
shared LSP server only during initialize, and deriving the version now
requires an async action call, so a handler can no longer just build
its settings dict in the constructor and call update_settings.
RuffLspService instead collects a settings provider from each handler
and resolves them all -- deep-merged so a linter's and a formatter's
contributions to the same table don't clobber each other -- right
before the server actually starts, regardless of which handler gets
there first.

Also:
- Add extend_ignore to the ruff lint handler, mirroring extend_select,
  so a preset can turn rules off without discarding a project's own
  [tool.ruff.lint] ignore list.
- Expand fine_python_lint's default rule selection (PLE, ASYNC, UP,
  DTZ, C4, PIE, PERF, FURB, SIM, RUF, FLY, INT, ICN, TID, PGH, LOG)
  now that a wrong target-version can no longer silently skew what
  these rules report.
- Bump ruff to 0.16.* and pyrefly to 1.2.*.
- Fix wm_server code using PEP 695 syntax (type X = Y, class Foo[T])
  that requires 3.12, which the presets' own requires-python (>=3.11)
  does not guarantee -- caught once ruff started targeting the range
  it actually declares instead of its py38 default.
Editing FineCode-managed code required restarting the whole workspace
server by hand to see the effect take hold, and there was no way for a
client to survive that restart without reconnecting manually.

Adds a four-rung recovery ladder, cheapest first, each covering what the
one before it doesn't (ADR-0073, ADR-0075):

- reload_action: re-imports the packages owning an action and its
  handlers, in every environment of every target project. Starts no
  process.
- restart_runner: replaces a project's extension runner processes, for
  edits to shared code no single action owns.
- reload_config: re-reads pyproject.toml/finecode.toml/presets and
  replaces the runners they configure -- the rung to reach for when
  unsure which kind of edit was made. ADR-0078 makes the target
  explicit: one project, or the whole workspace via allProjects, never
  defaulted into.
- restart_wm: replaces the workspace server process itself, for edits
  to FineCode's own code.

Recovery is a WM capability the client only projects or refuses
(ADR-0077): each command needs --shared-server, since recovering a
dedicated server the command itself started would report success while
changing nothing that outlives the command. `server/reset` is gone
(ADR-0076) -- it logged one line and returned {}, indistinguishable
from a reset that worked; `workspace/reloadConfig` with
allProjects/rescan replaces it.

A reload or restart is refused while a run is in flight in its target
project, since replacing runners kills whatever they're executing with
no way to tell the caller whether side effects happened;
killInFlightRuns overrides it by name (ADR-0079). Runs the WM started on
its own behalf are exempt -- re-derivable and awaited by nobody, they're
cancelled instead of blocking (ADR-0080). Both are tracked in the new
in_flight_runs registry.

Restarting a shared WM disconnects every other client, so ApiClient
grows reconnect-with-backoff and session re-attachment (ADR-0074): a
dropped connection retries with jittered exponential backoff inside the
WM's 30s disconnect timeout, then replays whatever session state the
surface needs -- add_dir, MCP tool-list invalidation, the LSP's
server_initialized gate -- through a single on_reattach hook shared by
first-connect and reconnect.

MCP exposes the four rungs as tools rather than actions, since an
action executing inside the runner it would replace can't complete.
Each tool's description states what it covers and what to reach for
next: staleness is never auto-detected (ADR-0075), so the description
is the only signal a caller has for which rung applies.

Also:
- schema_utils: describe nested dataclasses (e.g. Range, Position) as
  JSON object schemas instead of {}, needed to expose the new tool
  parameters correctly.
- next_step.for_runner_failure derives the `prepare-envs` command a
  failed recovery should be retried with from the runner's own failure
  state (PRD-0008 R11), instead of surfacing a raw import error.
- New docs/guides/wm-server-internals.md describes the WM layer stack
  and the recovery flow end to end.
Code actions were offered but had no way to be executed: the LSP
codeAction/resolve request had nothing behind it, and there was no
action to turn a get_lint_fixes/get_code_actions result into a write.
This adds the missing writing half of the code action pipeline.

- resolve_code_action: recovers a code action's edits from the
  (provider, action_id) pair round-tripped through the LSP `data`
  field, since the client only ever holds an opaque handle, not the
  edits themselves. Concurrent handlers each claim only the actions
  their own provider minted (lint_fixes_resolve_bridge_handler checks
  `provider == PROVIDER_ID`) and pass through untouched otherwise, so
  one resolve request fans out safely to every registered provider.

- apply_code_actions: applies a batch of selections (each pinned to
  the file version it was computed against) via a new text-edit
  algebra (_text_edit_algebra) that applies same-file edits
  back-to-front so ranges stay valid, detects overlaps, and reports a
  per-selection ApplyOutcome (applied / deferred / version_conflict /
  invalid_range / unresolved / unsupported_operation / write_failed /
  partially_applied) rather than failing the whole batch on one bad
  selection. Only TextEditOperation is executed; Create/Rename/Delete
  are refused wholesale until IFileEditor grows support for them.

- apply_lint_fixes / apply_lint_fixes_files: the fix-oriented entry
  points, layered on top of apply_code_actions so lint fixes and
  editor-offered code actions converge through the same apply and
  version-conflict handling instead of two parallel write paths.

The LSP code_actions endpoint now carries {provider, action_id,
file_path} in each CodeAction's `data` instead of a bare action_id,
and adds a cattrs structure hook to disambiguate the
CodeActionOperation union on the way back in (cattrs' default
disambiguator can't tell CreateFileOperation and DeleteFileOperation
apart from a shared file_path with only defaulted fields of their
own).

Also catches docs/reference/lsp-protocol.md up to the reloadConfig/
restartWm rename that shipped in fbdf049 but never touched this file.
fbdf049 (PRD-0008) described ADR-0080 in its commit message -- runs
the WM starts on its own behalf, re-derivable and awaited by nobody,
should be cancelled by a config reload instead of blocking it -- but
never wired the plumbing: InFlightRun had no cancellable field and
config_reload_service always refused on any run regardless of who
started it. This finishes that:

- InFlightRun.cancellable (domain.py) and in_flight_runs.track/
  blocking_runs now drop cancellable runs before checking whether a
  project is idle, only falling back to naming a blocking run when a
  user-started one is present alongside them.
- run_action, run_actions_in_running_project/_in_projects and
  WorkspaceExecutor.run_actions thread a cancellable flag through to
  the track() call, so a caller can opt a whole dispatch in; it is a
  property of the dispatch, not the action.
- shutdown_service flushes knowledge_service's pending fact writes
  before stopping runners, since a graceful shutdown has no reason to
  accept the throttled writer's usual crash-window loss.
- New integration tests exercise both halves: a lone cancellable run
  lets reloadConfig proceed, and a user run sharing the project with
  one still refuses and names only the run that matters
  (on-demand-extraction-plan D-8 AC7).

Also catches docs/wm-protocol.md and docs/cli.md up to the PRD-0008
rungs that fbdf049 shipped without documenting: reload_action and
restarts_runner's MCP exposure, allProjects/rescan addressing,
server/getInfo's pid and clients fields, and a recovery-ladder
quick-reference table in cli.md.

tests/integration/conftest.py exposes InProcClient.ws_context so
tests can seed in-flight runs and read workspace state back directly.
…project boundaries

Two independent fixes that landed together:

Ruff's CLI (`ruff check`) and LSP paths reported the same violation at
different positions: dropping the 1-based-to-0-based column shift put
every CLI-path range one character to the right of what the LSP path
and an editor agree on, so highlights covered the wrong span and
`payload.range` filtering missed matches across the two paths.
`_position_from_ruff` now does both the row and column shift in one
place. Also `_run_cli_fixes` was calling `ICommandRunner.run` with an
argv list where it takes a single shell string -- the CLI path only
ever worked against a stub; `shlex.join` fixes the real call.

Ruff reports fixability inline with each violation (`fix` non-null)
but the LSP protocol has no field for it, so `Diagnostic.fixable` is
threaded through from ruff's own `data` on the LSP path and from the
CLI's `fix` field on the other, with `None` kept distinct from `False`
so "no fix" isn't confused with "tool doesn't say." The CLI text
report marks fixable diagnostics with `[fixable]`.

LSP code actions carry no applicability, and ruff offers unsafe fixes
as quickfixes regardless of configuration, so a fix arriving over LSP
could not be told apart from a safe one. `_label_applicability` now
cross-references each LSP fix against a same-content `ruff check`
call (by code + position) to recover its real applicability, and
flags noqa-suppression fixes as `DISPLAY_ONLY` since they suppress
rather than fix.

Separately, `list_src_artifact_files_by_lang`'s Python and TOML
handlers walked their project directory with a plain `rglob`, which
does not stop at a nested project's root -- a workspace operation
scoped to the outer project silently pulled in the inner project's
files too, and an unscoped one processed them twice. New
`workspace_utils.nested_project_dirs` / `walk_project_files` prune the
walk at any nested project's root (using workspace project info that
was already available) and skip hidden directories/files (venvs,
`.git`, dotfile configs) by default.
The knowledge model engine (entity/fact model, query IR, interpreter,
memoization DAG) previously lived alongside FineCode's own schema and
rules. R20 requires the core to carry no language- or tool-specific
logic, but that was only a lint-enforced convention -- nothing stopped
rule code from ending up in the same distribution the WM imports to
run the DAG.

Splitting it into its own package makes R20 a packaging fact instead:
finecode_knowledge has zero dependencies and no schema of its own:
schema, providers and rules are declared by a separate distribution
(fine_knowledge) and handed in as a SchemaRegistry via
set_default_registry. The WM can import the engine without any rule
code entering its process, because the package holding rules is never
installed in its environment.
resource_uri.py already flagged the bug in a NOTE: a relative
file:// URI reaching an ER resolves against that ER's own project
directory, not the user's terminal directory, so the same payload
silently names a different file per project a run fans out to. The
CLI is the last process that still knows the user's directory, so
that is where expansion has to happen.

- resource_uri.py gains absolutize_resource_uri, resource_location_to_uri
  and is_relative_file_uri, splitting out _parse_file_uri_path so the
  netloc/path reassembly ("file://relative/path" splits in two under
  urlparse) is shared instead of duplicated.
- New cli_app/payload_uris.py walks an action payload against the
  fields its own payload schema (actions/getPayloadSchemas) marks
  format: "uri", and rewrites only those to absolute file:// URIs —
  a plain path or a relative URI is accepted, but only on a field a
  schema vouches for as a resource.
- run_cmd.run_actions calls it before dispatch and now fails the run
  if a relative file:// URI survives, naming the field and whether a
  schema was even available, rather than sending something an ER
  would resolve wrong.
- mcp_server: stop advertising a "project" tool argument for
  workspace-scoped actions, since the WM always resolves those to
  the workspace root itself and rejects an explicit one; and replace
  the top-level try/except KeyboardInterrupt with
  contextlib.suppress.
Project paths on finecode/runActionInWorkspace come from a handler's
payload, typically a caller-supplied URI some ER turned into a path,
so the WM has never vetted them. An unknown path used to reach a bare
dict lookup deep in the fan-out (proxy_utils.run_actions_in_projects
or the actions_by_project dict in er_dispatch), and the caller was
handed a KeyError whose entire message was the repr of a PosixPath:
it named the path but never said what it had failed to match.

- er_dispatch and proxy_utils now check requested paths against
  ws_context.ws_projects up front and raise errors.ProjectError
  naming the path, the requesting runner, and the closest known
  projects by shared path segments (capped at 3, "and N more") so a
  hundred-project workspace doesn't bury the answer.
- partial_results_service rejects an explicit project path passed
  alongside a workspace-scoped action, mirroring the non-streaming
  actions/run guard in _helpers.py, instead of silently narrowing a
  workspace-wide action to one runner.
- run_dispatch_bridge documents the new ProjectError case.
AsyncProcess only ever buffered a child's stdout/stderr and handed it
over on wait_for_end(), so any handler wanting live output had no way
to get it -- get_output()/get_error_output() carried "TODO: live
output?" comments marking exactly this gap. stdout_lines()/
stderr_lines() close it without disturbing the ~20 existing handlers
that still read output the old way.

- ICommandRunner gains stdout_lines()/stderr_lines() on IAsyncProcess
  only -- the sync flavour has no way to interleave reads with
  anything else.
- command_runner.py: _LineStream drains a stream unconditionally from
  spawn (an unread pipe fills its OS buffer and blocks the child
  forever), accumulates until a subscriber shows up, then switches to
  a queue and stops accumulating. A stream that hits an unrecoverable
  error (a line past the reader's size limit, or non-UTF-8 output)
  keeps draining and discarding rather than stopping, so one bad
  stream cannot hang the whole process. The per-line limit is raised
  well past asyncio's 64 KiB default, since an overrun there discards
  the buffered bytes before raising rather than just reporting it.
- CommandRunner now keeps strong references to its release-when-done
  tasks. The event loop only holds weak ones, so a task with nothing
  referencing it could be collected before the child exits --
  permanently losing a semaphore slot and eventually wedging the
  runner at its concurrency cap.
- Test doubles across the ruff, uv and git presets/extensions gained
  stdout_lines()/stderr_lines() to keep satisfying the IAsyncProcess
  protocol.
A shared server started by run --shared-server still exits 30s after
the last client disconnects — the flag amortizes nothing if no
process holds the server up between calls, so config loading and
runner startup were still paid on every command in the devcontainer.

- start-wm-server gains --keep-alive (disables both the no-client and
  disconnect auto-stop timers; server/shutdown still stops it) and
  --detach (start it in the background, no-op if one is already
  listening). Keep-alive is passed explicitly and never read from the
  environment, so it can't leak onto the dedicated per-command
  servers, which must keep auto-stopping.
- wm_lifecycle.ensure_running forwards keep_alive/disconnect_timeout/
  wal_enabled to the spawned server and starts it in its own session,
  so it survives the signals of whichever short-lived client or
  script happened to start it.
- .devcontainer/start-wm-server.sh runs start-wm-server --detach
  --keep-alive from postStartCommand, gated on FINECODE_WM_AUTOSTART
  (on by default in .env.example) and skipped when the dev_workspace
  venv doesn't exist yet.
A run_agent_task action lets FineCode delegate a task to an AI coding
agent without a caller committing to a specific backend: the fine_agent
preset registers the action slot alone, and a backend extension supplies
the one handler. fine_system_claude_code is renamed to
fine_agent_claude_code to sit alongside the new fine_agent_pi extension
under that naming, and ICommandRunner gains process-group teardown so a
handler can actually stop a hung agent run -- signalling the shell alone
leaves the tools it spawned running.

- presets/fine_agent: RunAgentTaskAction, RunAgentTaskRunPayload/Result,
  and AgentRunUsage carrying only what a backend actually reported,
  never a derived total or a converted currency.
- extensions/fine_agent_claude_code (renamed from fine_system_claude_code):
  ClaudeCodeAgentHandler drives the Claude Code CLI over its
  stream-json protocol.
- extensions/fine_agent_pi (new): PiAgentHandler drives pi.dev over its
  own RPC-like wire format (pi_rpc.py), which is not JSON-RPC and has
  several surprises documented there (agent_settled vs agent_end,
  doubled usage figures across message_end/turn_end).
- finecode_extension_runner/impls/command_runner.py: run(...,
  new_process_group=True) starts the child in its own session;
  is_alive()/terminate()/kill() signal the whole group so a handler's
  escalation ladder (SIGTERM, wait, SIGKILL) reaches an agent's own
  tool-call subprocesses, not just the wrapping shell.
- docs/reference/services.md documents the new stop-a-process pattern.
fine_git could only push tags, so nothing in the workspace could read a
project's git state or discard local changes through an action -- a
caller had to shell out itself. get_git_status/get_git_diff report both
porcelain columns and both diff projections (patch and parsed lines) in
one answer, since callers want different slices of the same state
rather than each parsing it themselves; restore_git_files requires an
explicit path list, since it destroys uncommitted work and "restore
everything" has no place being one flag away.

- get_git_status_action.py / git_get_git_status_handler.py: reports
  repo_root=None for a project outside a git repo as a result, not an
  error.
- get_git_diff_action.py / git_get_git_diff_handler.py: source picks
  worktree/staged/worktree_and_head; context_lines=0 yields hunks with
  only changed lines.
- restore_git_files_action.py / git_restore_git_files_handler.py: a
  path outside the project directory is refused and reported in
  skipped rather than restored; deleting an untracked path needs
  remove_untracked=True since it is unrecoverable.
- tests/conftest.py: FakeCommandResult/FakeCommandRunner moved out of
  test_create_git_tag_handler.py so the three new handlers' tests can
  share them.
ICommandRunner.run()/run_sync() took a single command string and
spawned it through a shell. Every caller therefore had to quote for two
different parsers (shlex on POSIX, cmd.exe on Windows), and arguments a
shell reinterprets (spaces, `%`, `"`, `&`, `>` in a version spec) either
broke or were a latent injection vector. The API now takes an argv list
and hands each element to the program as exactly one argument on any OS.

- icommandrunner: add the `Argv` type and `check_argv`, which rejects a
  bare `str` (it satisfies Sequence[str], so the type checker alone
  would let it through and it would exec character by character) and
  non-str elements such as a `Path`. Add `CommandNotLaunchableError`
  with `UnsafeBatchArgumentError` and `UnlaunchableProgramError`.
- command_runner: spawn via create_subprocess_exec. On Windows resolve
  bare names through PATHEXT (CreateProcess only appends `.exe`), accept
  only .exe/.com/.cmd/.bat, and refuse an argument a batch file would
  reinterpret rather than trying to escape it.
- An unstartable program now raises OSError / CommandNotLaunchableError
  instead of returning a nonzero exit. Agent, installer and pi package
  handlers catch it and report a structured FAILED result through the
  new backend_support.spawn_error helper.
- Convert all handlers (pip, uv, mypy, pyrefly, pytest, ruff, mkdocs,
  git, package_info, agent backends) from `shlex.join`/f-string commands
  to lists, dropping the per-handler `_quote_arg` workarounds.
- Document the contract in docs/reference/services.md, including the
  is_alive() vs get_exit_code() note, which no longer involves a shell.
- Bump finecode_extension_api to ~=0.5.0a0 in every dependent package,
  since the ICommandRunner signature change is breaking.
…s (ADR-0100)

The startup semaphore covered only spawn until the RPC channel
connected, but the rest of the start burst (initialize, runner info,
preset resolution, updateConfig) is just as CPU/import heavy, so more
ERs were mid-start than the cap intended. Widening the hold to RUNNING
alone would deadlock: an INITIALIZING ER can call back into the WM for
something that needs another runner to start, and that runner would be
queued behind the slot its caller holds.

- _start_runner now owns the slot (`_StartupSlot`, idempotent release)
  from just before spawn until RUNNING. The JSON-RPC client is attached
  before the acquire so a shutdown sweep can force-kill a queued runner
  (ADR-0097).
- Back-channel methods are classified as yielding (may wait on a
  runner, the process budget, knowledge extraction or a human) or
  neutral. A yielding call releases the slot once before its handler
  runs. An unclassified method fails open: yield and log a warning,
  so a missed classification widens the gate rather than deadlocking.
- Preset resolution raises DevWorkspaceRunnerNotConnectedError instead
  of an AttributeError when the dev_workspace runner is still queued,
  and never waits on it while the caller may hold a slot.
- Log spawn-to-output/port/connected timings on slow starts and the
  spawn age on a failed start.
- Add test_no_shell_command_strings, gating shell spawns and string
  commands out of ICommandRunner callers (ADR-0099), and extend the
  startup-concurrency tests to cover yielding and classification.
- Document the widened hold and the known small-host overshoot from
  ADR-0094's stall escape in wm-server-internals.md.
The startup-timeout snapshot read /proc directly, so it returned an
empty string off Linux and a failed start on Windows or macOS said
nothing about what the spawned server was doing at the deadline.

- Replace describe_process_group with describe_spawned_processes,
  built on psutil. On POSIX the server is spawned with
  start_new_session=True, so its pid is its process group and every
  member is listed, including children reparented after their parent
  exited. Windows has no process group, so the tree rooted at the
  spawned pid is listed by parent pid, falling back to processes whose
  ppid is the root when it already died. Windows lines omit state=
  (psutil reports `running` for nearly every process) and add ppid and
  thread count.
- An unavailable snapshot now says why ("process snapshot unavailable:
  <reason>") instead of a fixed string, and never raises so it cannot
  mask the start error it explains.
- Declare psutil as a finecode_jsonrpc dependency; it was previously
  only a dev/test dependency of the root project.
- Cover the per-platform behaviour in test_proc_snapshot and the
  ServerFailedToStart message in test_startup_timeline.
- Update wm-server-internals.md; the Job Object replacement that would
  make Windows membership a spawn-side record is tracked in issue 43.
The CLI could start a shared workspace server (start-wm-server
--detach --keep-alive) and recover it (restart-wm and friends), but had
no counterpart to stop it. A keep-alive server outlives its clients by
design, so the only way to end one was to kill the process by hand.

- Add `stop-wm` alongside the other recovery commands. Like them it
  requires --shared-server: stopping a private per-command server would
  report success while leaving the workspace an editor or agent uses
  untouched.
- Refuse with RecoveryFailed when no server is running, naming how to
  start one, rather than reporting a stop that did nothing.
- Send server/shutdown through the new ApiClient.shutdown and tolerate
  the server closing the socket before the response is read, as
  replace_running_server already does; the wait that follows verifies
  the stop.
- Add wm_lifecycle.wait_until_stopped, which polls running_port in a
  thread (its synchronous probe blocks for the full timeout on a
  filtered port) and fails with a timeout naming the port.
- Document the command in docs/cli.md and update "all four" to "all
  five".
When a handler's TaskGroup crashed with an unexpected exception,
_classify_exception_group returned str(eg), i.e. "unhandled errors in a
TaskGroup (1 sub-exception)". That text is what every wrapper up to the
caller and the streamed logs embed, so the actual error (a cp1252
UnicodeDecodeError on the Windows CI, issues #46/#47) was visible only
in the ER's own log.

- Flatten nested exception groups to their leaves before classifying,
  deduplicating by identity so an exception wrapped at several levels
  is listed once while distinct failures with equal text all show.
- Summarize an unexpected leaf as "Type: message", falling back to repr
  when the message is empty. Known failures keep their own message
  verbatim, as a deeper layer already summarized them.
- Treat asyncio.CancelledError as a benign cancellation and drop empty
  messages, so a message-less cancelled sibling neither counts as a
  crash nor leaves a trailing "; " in the summary.
- Add test_exception_group_summary, including an end-to-end handler
  whose own TaskGroup raises.
FileManager and LspService opened files with the locale default
encoding. On Windows that is cp1252, so a source file containing
characters outside it (e.g. U+201D or an emoji) raised
UnicodeDecodeError on read, and writes produced non-UTF-8 bytes. Linux
and macOS hide this behind UTF-8 locales, which is why #46 and #47
showed up only on the Windows CI.

- Pass encoding="utf-8" in FileManager.get_content/save_file and in the
  LspService file-open read, and document the UTF-8 contract on
  IFileManager.
- Add regression tests using text outside cp1252, so a locale-decoded
  read raises rather than silently producing mojibake.
- Record the convention in developing-finecode.md, with the env vars
  that reproduce a non-UTF-8 locale. PLW1514 is preview-only and misses
  method references passed as callables, so enforcement stays manual.
A plain `run` started every interpreter instance of a matrixed env while
the run itself only executed the selected subset, so unselected children
could fail to start (or be repaired) for no reason. Separately, an ER
that died before publishing its port surfaced as a raw start error even
though a reinstall usually fixes it.

- Thread the run's env/interpreter selection into the payload-schema
  fetch (`runOptions`, honoured only with `startRunners`). The CLI builds
  one dict that feeds both the fetch and the run, so they cannot drift.
  Selection is computed only for projects with a matrixed action and
  selectors are never validated in the fetch, so a selector valid in a
  sibling project cannot fail it.
- Repair a runner that crashed before its port (ServerExitedBeforePort
  in the __cause__ chain) the same way as NO_VENV: install, then
  restart. Only at the run gate and dispatch start, never during
  metadata resolution, and never for timeouts, which are load problems.
- Rename repair_no_venv_env to repair_env and serialize repairs per
  (project, env) so two callers never install into one venv at once.
- Pass the env's configured interpreter to the create_envs step in
  install_env_for_project so a repair rebuilds the venv with the right
  Python.
Both suites failed only on the Windows CI, for reasons in the tests
rather than the code under test.

- Command runner streaming tests wrote through text-mode stdout/stderr,
  where Windows translates `\n` to `\r\n`. A literal `\r\n` therefore
  became `\r\r\n` and the buffered/CRLF assertions saw different bytes.
  Write bytes via `sys.stdout.buffer` / `sys.stderr.buffer` instead so
  the child emits exactly what the test asserts on.
- fine_git handler tests used a POSIX literal `/repo` as the fake
  toplevel. It has no drive on Windows, so it is not absolute and
  cannot become a `file://` URI. Add a `repo_root` fixture backed by
  `tmp_path` and a `toplevel_result` helper that prints it with forward
  slashes, as Git for Windows does, and derive expected URIs from it.
The setuptools_scm handler makes setuptools_scm write the configured
`version_file`, and that file is then left in setuptools_scm's own
layout, so the format check flags it after every version query.

- Add FormatSetuptoolsScmVersionFileHandler, registered after the scm
  handler. It runs `format_file` on the written version file and returns
  the version unchanged. Which formatter runs stays `format_file`'s
  dispatch decision; a file no formatter covers is reported as coverage,
  not absorbed, while a formatter failure fails the run.
- Read the content from disk rather than through the file editor:
  setuptools_scm writes past the editor, so an open buffer would hold the
  previous version and saving it would overwrite the fresh write.
- Pass `force_write_version_files=True` so the file is rewritten on every
  query and there is always a fresh file to format.
- Move config loading and version-file path resolution into
  `_scm_config` so both handlers share them.
- Register the handler in the root and extension-runner pyproject, add
  the `fine_format` dependency, document it in the actions reference,
  and add tests for both handlers.
The cold `prepare-envs` install alone can run for about two hours on the
slower matrix legs (uv rebuilds the local packages in every env), so the
120-minute job timeout cancelled runs before the check steps could
report anything. A single job-wide limit also hides which step hung.

- Raise the job timeout to 240 minutes and give the install (80), inspect
  (30) and audit (30) steps their own limits, so a hang in one step fails
  that step instead of eating the whole job budget.
- Trim the uv cache whenever the install step ran, not only when it
  succeeded, and save it once the trim succeeded. A timed-out cold
  install has already built most of the wheels; discarding them made the
  next run pay the same two hours again.
A project's action set is only complete once its presets resolve, which
needs a running dev_workspace ER — so eagerly starting runners for every
project on add_dir (or for every --project on the CLI) forced all of
them through that cost even when a run only touches one. Push
resolution to the entry points that actually need a project's actions
or preset-dependent config, and let everything else stay CONFIG_VALID
until asked for.

- Add project_resolution_service: ensure_projects_resolved /
  ensure_all_projects_resolved gate every read of a project's actions,
  waiting on in-flight resolution, remembering attributable failures
  until the project's config is reloaded, and reporting unresolved
  projects instead of skipping them silently.
- Add resolve_hosting_projects for root-first resolution: a run whose
  actions are all workspace-scoped and root-hosted resolves only the
  root; anything else resolves every project and fails loudly on a
  failed sibling rather than dropping it from the result.
- add_dir now always starts with start_runners=False across the CLI,
  LSP and MCP clients; runners are started lazily through the
  resolution gate instead.
- actions/list and the wm_client ActionListing carry unresolvedProjects
  so callers (CLI run, MCP tool listing) can report a partial listing
  instead of failing or silently omitting projects.
- Extract action_lookup.py (find_action_by_source, project_exposes_action)
  out of _helpers.py so project_resolution_service can use it without a
  layering cycle.
- Move handler-config-override application into collect_actions, applied
  once per project at collection time, replacing the add_dir-time
  re-push to already-running runners that lazy start now makes
  unreliable.
Continues the Windows CI stabilization work: the runner itself had two
behaviors that are wrong on Windows, and the test suite had hardcoded
POSIX absolute paths that cannot exist there.

- Decode process output with errors="replace" instead of failing the
  whole stream on the first undecodable byte. No consumer derives a
  run's verdict from raw stdout bytes, so a child writing in the
  console encoding (cp1252 on a Windows runner) should cost a few
  U+FFFD characters, not the run.
- Add a psutil-backed process tree kill for Windows, where there are
  no process groups: a direct signal only stops the child and orphans
  everything it spawned, which was leaving grandchildren running with
  the temp dir pinned open (WinError 32 on cleanup). `terminate()` and
  `kill()` both route through it there; add psutil as a win32-only
  dependency and a Windows-only teardown test file.
- Add `nonexistent_abs_path()`, an in-memory identity path that spells
  correctly as absolute on whichever platform the test runs on, and use
  it in place of literal `/tmp/...` and `/venv`-style paths across the
  pip/uv/ruff/pyrefly/fine_agent/check_imports/inlay_hints test suites.
- Rework the resource_uri tests off a hardcoded `/ws` literal onto a
  tmp_path-backed fixture, and make the relative-URI assertions
  platform-aware where POSIX and Windows genuinely disagree on what a
  leading slash means.
- Replace an elapsed-time assertion in the LSP service's watched-file
  sweep test with a direct check of republish ordering, removing a
  timing race that flaked under Windows CI's slower scheduling.
The 80-minute budget set for the install step still isn't enough on
the slower matrix legs, where uv's cold rebuild of local packages
across all envs can run close to two hours; give it more headroom so
a legitimately slow but healthy install isn't cancelled mid-run.
The .ignore re-includes the gitignored nested finecode_internal_docs/
for ripgrep-based search (VS Code, Claude Code); without it the moved
docs are unsearchable from the root. developing-finecode.md gains the
matching clone-command paragraph (issue 59, plan step A8).
A run-scoped lease made every action run hold slots for its whole
duration, including runs that spawn nothing and parents that only wait
on children. Nested runs then needed a "nested, always at least one"
escape and per-project shares (RunBudget, prepare-envs `W // N`) to keep
from deadlocking against their own ancestors.

Move the lease to where the work happens. ProcessSlots now takes a WM
lease around each bounded subprocess (ICommandRunner.run) and
IProcessExecutor task, so a run that spawns nothing holds nothing and
parents never hold slots their children wait on. Every lease waits;
granted stays <= work_cap + 1 (the +1 is the stall escape). A backend
failure degrades to the local gate, since a throttle never refuses work.

- Add IWorkSlots.acquire() for CPU-heavy work sent to long-lived local
  servers (LSP services: ruff, pyrefly, tombi). Operations that can wait
  on a slot raise WorkSlotScopeError inside the scope.
- LspService bounds its requests with the explicit scope.
- Drop the run-level lease/release, RunBudget, per-ER gate resizing
  (update_process_budget, target_for_runner) and prepare-envs' per-project
  budget, all made redundant.
- Cancelled acquires return the local slot at once and release any
  pending WM grant from a strongly referenced cleanup task, so a
  cancellation cannot leak a lease.
- Update docs and tests to the per-unit model.
CI jobs that die under memory pressure leave nothing behind but the job
log, and the WM log goes with the lost runner. Operators also had no way
to ask a live workspace server what it was doing or what it cost, so
diagnosing ER fan-out, slot starvation and swap thrash meant guessing.

- Add the read-only `server/getResourceUsage` WM request: runner,
  project, startup-slot and work-slot counts, in-flight runs, host
  memory/swap/cgroup/PSI, event-loop lag and peaks since server start.
  It never waits: no lock, no lease, no ER request, no runner start.
- Add an optional per-process footprint walk (VmRSS + VmSwap per WM and
  ER tree, untracked ERs as their own rows) behind `includeProcesses`,
  run off-loop one walk at a time. Adds psutil for non-Linux hosts.
- Add `finecode resource-usage --shared-server [--json] [--watch]
  [--processes]` and an MCP `get_resource_usage` tool.
- Add `--resource-usage[=SEC]` / `--no-resource-usage` to `run` and
  `prepare-envs`: one flushed `[resources]` stderr line per interval
  plus a peaks summary, on by default when CI is set.
- Centralise runner counting in runner_counts.py so the lag monitor and
  the snapshot share one status filter; record peaks only where a count
  rises, with hooks that never raise into the code they observe.
- Keep a lag window in the event-loop monitor so a stall is visible to
  the next poll, and expose ProcessBudget.snapshot() with peaks.
- Document the protocol method, CLI options and snapshot definitions.
A filtered run created venvs for every non-matrix env but installed
into only the named ones, so `--env=x` still paid for creating (and,
with --recreate, wiping) envs it never touched. `--recreate` with a
filter also wiped every subproject dev_workspace venv, the bootstrap
every later step runs on.

- Replace compute_create_set/compute_install_set with a single
  compute_prepare_set, so both steps cover exactly the selected envs
  and cannot drift apart.
- Extract build_install_envs_params, sharing that set with the create
  step; dev_workspace stays excluded from both.
- Wipe subproject dev_workspace venvs on --recreate only when no --env
  filter is given or it names dev_workspace. The bootstrap validity
  check and rebuild still always run.
- `--interpreter` alone keeps every non-matrix env selected.
- Update CLI docs, the preparing-environments guide and the handler
  docstring. The trade-off is documented: an unselected venv that is
  missing or broken stays so until the next unfiltered run or an
  on-demand repair.
…-env (ADR-0103)

`--env` and `--interpreter` both existed on `run` and `prepare-envs`
and intersected in ways callers could not predict: a base named by
`--env` silently ignored its `default_interpreters` policy, and
`--interpreter` on prepare-envs widened or narrowed every matrix base
at once. Each command now has one selector, matching what it operates
on: `run` fans out over interpreters, `prepare-envs` over envs.

- `run` selects by `--interpreter` only; `all` means the full axis and
  ignores `default_interpreters`. `--env` is removed from `run`.
- `prepare-envs` selects by `--env` only, in four forms: `<base>` (the
  base's default subset), `<base>@<impl>-<version>` (one child),
  `<base>@all` (every child) and a non-matrix env name. `--interpreter`
  is removed from it.
- With any `--env` given, a base not named contributes nothing; without
  one, every base applies its own default policy. Each base's policy
  selects its own children only.
- Rework env_selection around resolve_run_selection, which yields
  concrete selected envs, and pass them down through matrix_runner and
  matrix_streaming via selected_variants so both fan-out paths filter
  identically.
- Update CLI docs, the preparing-environments guide, WM internals and
  protocol docs, and the unit tests.
…ends

The reporter advanced next_tick after each poll finished, in seven
separate exit paths. A slow poll therefore pushed the whole grid back,
so a late-finishing poll shifted every later sample off its interval
and a poll that started late could be blamed on the server with a false
"no answer" line. A false line trains operators to ignore the one
signal that means the server is actually stalled.

- Compute next_tick once, at poll start, from max(next_tick, start) so
  a late start still earns a full interval and a stall is reported once
  per missed tick before sampling resumes on the grid.
- Drop the per-path `next_tick += interval_sec` increments.
- Add virtual-time tests (a selector loop that jumps the clock to the
  next timer) covering prompt answers, a slow poll with missed ticks,
  and a late-started poll.
- Harden neighbouring tests: bound the "no answer" wait with a timeout,
  give _FakeWm a handler lambda, and widen a sleep that raced the tick.
A CI job that dies under memory pressure surfaces as an ER boot stall or
a generic run failure, and nothing in the client output says the host was
thrashing. Operators had to correlate the failure with swap graphs after
the fact, or guess.

- Add a host memory-pressure predicate in host_pressure: PSI `memory
  full avg10` >= 10%, or available memory <= 5% of total with swap
  exhausted or absent. Host-wide /proc figures only. It reports
  "unevaluable" instead of "calm" when neither PSI nor memory figures
  are readable.
- Expose it as `host.memoryPressure` and a `hostPsiMemoryFullMax` peak
  in the resource-usage snapshot, and show `psi` in the `[resources]`
  line.
- Have the CLI reporter print one warning when a pressure episode
  starts and one when it clears. An episode needs two consecutive calm
  polls to close, so one calm sample does not end it. A footprint line
  follows from a single per-episode process walk, plus one for the
  final peaks line under a 3 s budget so a slow walk never delays it.
- Route ActionRunFailed / StartingEnvironmentsFailed handling in the
  streaming handlers through one client_message helper. It logs the
  failure with host state as structured extras and appends a
  `[host under memory pressure: ...]` note to the client message only
  while pressure is active.
- Update the CLI docs, WM protocol and WM internals guide.
Projects had no way to downgrade or silence individual pyrefly error
kinds (e.g. implicit-any-type-argument) through FineCode. The setting
is per server, not per handler: the pyrefly handlers share one LSP
server per runner, so a handler-level surface would let them disagree
and apply the setting late.

- Add the self-bound PyreflyConfig service with an `errors` table
  mapping error kind to error/warn/info/ignore. CLI mode turns it into
  --<severity>= flags (plus --min-severity when needed); LSP mode
  writes a generated pyrefly.toml in the runner cache dir, validates
  it with `pyrefly dump-config`, and passes it as `configPath`.
- Fail early, in both modes, when a project pyrefly config exists
  alongside a non-empty `errors`: pyrefly would silently ignore one of
  the two.
- Push settings registered after the LSP server started via
  didChangeConfiguration instead of losing them.
- Map pyrefly's reported severity to the diagnostic severity instead
  of always using ERROR, and include pyrefly stderr when its output is
  not JSON.
- Report any malformed service config entry as ActionFailedException
  naming the service: list- and dict-typed fields raise a cattrs
  error that is not a ClassValidationError and was escaping unwrapped.
- Require pyrefly >=1.3.0 and document the service and its blast
  radius in the extensions and services references.
install_deps_in_env ran `uv --no-config pip install`, so uv never read
the temporary pyproject.toml dump the handler wrote for it. Each env
install still paid for fetching the project's raw config and rendering
it (~220 ms per env), and the extra fetch is what timed out
getExtraSelection under load.

- Run uv directly in the project dir; the dependency specs on the
  command line are its complete input.
- Remove the action_runner and project_info_provider dependencies from
  the handler, and the "inspect the dump" hint from its error message.
- Note next to --no-config why no dump is written here, while
  create_env's `uv venv` still reads one.
- Replace the dump tests with one asserting no config is fetched and no
  dump action is dispatched, and trim the _uv_common docstring that
  named install as a dump consumer.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant