Skip to content

Add YYLO Benchmark to AI & LLM Testing - #239

Open
InsightFactoryAPP wants to merge 1 commit into
TheJambo:masterfrom
InsightFactoryAPP:add-yylo-benchmark
Open

InsightFactoryAPP wants to merge 1 commit into
TheJambo:masterfrom
InsightFactoryAPP:add-yylo-benchmark

Conversation

@InsightFactoryAPP

Copy link
Copy Markdown

Hi maintainers! Thanks for maintaining this valuable list.

I'd like to propose adding YYLO Benchmark to the AI & LLM Testing section.

What is YYLO Benchmark?

YYLO Benchmark is a flexible, isolated evaluation layer for task prompts and project-owned Workflow Runner YAML. It normalizes every case into one attempt contract, runs each candidate in a private fresh-repository workspace (dedicated repo plus private home/temp/cache roots behind a selective filesystem sandbox), and evaluates the retained evidence with ordered deterministic and/or LLM judge profiles. Results chain receipts, manifests, and hash-verified terminals into immutable records, so regrade/rejudge consume retained evidence instead of re-running candidates.

Why this section?

The section's focus is exactly this noun: testing frameworks for AI systems (promptfoo, Tenro, nika, crilio). YYLO Benchmark occupies the agent-run side — the agent's executed attempt is the software under test, evaluated deterministically and by LLM judges with retained, tamper-evident evidence. It complements prompt-output regression tools like crilio and workflow snapshot engines like nika.

Contribution notes

  • Single addition at the end of the ### AI & LLM Testing section, one-line format matching the existing entries, description ends with a full stop, no trailing whitespace.
  • Searched the list first — no duplicate (no existing YYLO entry).

Disclosure

I'm part of the team that builds YYLO. This PR was prepared with an AI coding agent and reviewed by me before submission.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant