Skip to content

feat: add LLMRuntime for prompt and model routing - #1778

Draft
paul-paliychuk wants to merge 4 commits into
mainfrom
paul-paliychuk/configurable-prompts
Draft

paul-paliychuk wants to merge 4 commits into
mainfrom
paul-paliychuk/configurable-prompts

Conversation

@paul-paliychuk

Copy link
Copy Markdown
Collaborator

Summary

  • Add opt-in Graphiti(..., llm_runtime=) that couples one LLMClient transport with a required default LLMModel, optional PromptRoutes, and optional LLMPromptOverrides.
  • Users can override prompt text and send specific prompts to a different model id on the same provider. Response schemas stay in the builtin registry and are not user-configurable.
  • Per-prompt model selection passes model= / small_model= into generate_response for that call. The transport is never cloned or mutated. Omitting the new kwargs keeps today's self.model / self.small_model behavior.

Motivation / context

Ingest already builds prompts and calls the LLM in many places. Customers want (1) a few prompt rewrites, (2) cheaper/faster models on 1–2 prompts, and (3) existing provider subclasses — without a second Graphiti instance or a new client hierarchy. The legacy llm_client= / prompt_library= path stays unchanged and is mutually exclusive with llm_runtime.

Impact

  • Public API: LLMRuntime, LLMModel, PromptRoutes, LLMPromptOverrides. Nested dataclasses so unknown prompt names are constructor/type errors. No Graphiti nicknames — bind LLMModel instances to local variables.
  • Compatibility: generate_response(model=, small_model=) is keyword-only and optional. v1 is multi-model on one transport; a Claude id on an OpenAI client is unsupported. GLiNER2 native extraction cannot be routed per prompt.
  • Resolution: model for prompt P = exact route, else group route, else default. Text = model overrides, else general overrides, else builtin.

Testing

  • Unit tests for runtime routing/overrides, complete_prompt facade, PromptName, and per-call model= on OpenAI / generic / Anthropic / Gemini / base client.
  • make lint (ruff + pyright) passed. Targeted pytest: 122 passed on the files above.
  • Not run: full make test against Neo4j, live provider integration tests.

Risks / follow-ups

  • Custom LLMClient subclasses with the old _generate_response signature still work on the legacy path; using them inside LLMRuntime requires accepting the new kwargs (or **kwargs).
  • Groq / Anthropic / generic OpenAI still ignore ModelSize.small (unchanged). Traces still record model.size, not the resolved model id.
  • Reserved LLMModel fields (temperature, structured output, …) are not wired yet. Cross-provider routing is out of scope for v1.

Test plan

  • Construct Graphiti(..., llm_runtime=runtime) and confirm llm_client / prompt_library together raise
  • Override one prompt builder and confirm only that prompt changes
  • Route extract_nodes.extract_attributes to a second model id; confirm other prompts stay on the default
  • Omit small_id on the default model and confirm ModelSize.small still uses the transport small_model
  • Confirm existing Graphiti(llm_client=..., prompt_library=...) callers are unchanged

Made with Cursor

paul-paliychuk and others added 4 commits August 7, 2026 11:19
Allow Graphiti callers to inject prompt libraries (or partial overrides via
create_prompt_library) so prompt selection is per-client instead of process-global.

Co-authored-by: Cursor <cursoragent@cursor.com>
…facade

Replace Protocol/TypedDict prompt wrappers with ABC groups returning ChatPrompt,
add fixed PromptSpec schemas, migrate maintenance LLM stitches through
GraphitiClients.complete_prompt, and add opt-in PromptBoundLLM multi-model
routing on a single provider transport.

Co-authored-by: Cursor <cursoragent@cursor.com>
Replace PromptBoundLLM with an opt-in runtime that routes prompt text and
model ids on one transport via generate_response(model=, small_model=),
without cloning or mutating the client.

Co-authored-by: Cursor <cursoragent@cursor.com>

This branch was successfully deployed

1 active deployment
development — 728b8d3b Deployed Aug 18, 2026 by paul-paliychuk via live-mcp-tests #172
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant