A conversational AI agent ("Kit") designed to guide high school robotics students through weekly metacognitive reflection on their team's regulatory processes. Grounded in Self-Regulated Learning (SRL; Winne & Hadwin, 1998) and Socially-Shared Regulated Learning (SSRL; JΓ€rvelΓ€ & Hadwin, 2013), the agent helps students reflect on how their team set goals, planned, monitored progress, and adapted β in a short session after each team meeting.
- Docker and Docker Compose
- Git
git clone <repo-url>
cd AgenticRoboticsEvaluator/infra
docker compose up --build
# Open the app (admin user is auto-created on first run)
open http://localhost:3000Login: admin / admin123
| Service | Port | URL |
|---|---|---|
| Frontend | 3000 | http://localhost:3000 |
| Backend API | 8000 | http://localhost:8000 |
| PostgreSQL | 5433 | localhost:5433 |
- Quick Start
- Project Overview
- Current Status
- Architecture
- Tech Stack
- Project Structure
- Key Components Explained
- What's Next
- Development Guide
- Documentation
High school robotics students benefit from reflecting on their teamwork, but coaches have limited time for 1:1 conversations. Students need a supportive "near-peer" they can talk to weekly after team meetings β one that focuses on how the team worked together, not just the robot.
A chat-based AI agent that:
- Guides students through a five-question SSRL reflection protocol (task/goal understanding, strategy & plan, monitoring & adaptation, shared motivation & emotion regulation, evaluation), each with signal-based follow-up branching
- Uses Socratic questioning with a coordination pivot when a student describes only technical work with no teammates in the picture
- Carries a loop-closure mechanism: each session opens by checking in on the prior session's action item, and closes by capturing one concrete commitment for next time
- Detects and responds to safety disclosures with a fixed, deterministic reply and a dashboard-visible flag, without ever changing the session's stage
- Enforces a single-question, near-peer voice with hard guardrails (no em dashes, one question mark, under 45 words) on every generated turn
- Near-peer tone: Like a slightly older student, not a teacher or coach
- Regulation-focused: The robot is context; how the team regulates their work is the subject
- Hybrid neurosymbolic design: an LLM classifies each turn (Layer 1), a deterministic Python policy decides what happens next (Layer 2), and a second LLM call phrases the reply (Layer 3) β the LLM never decides whether to advance, probe, or close
- No ground truth: the agent never observes the actual team meeting, only what the student reports β it never asserts a verdict on how the team did
- Privacy-conscious: Minimal data collection, clear boundaries, cross-student data isolation
The core system is fully functional with LLM integration and a dashboard UI.
| Layer | Status | Description |
|---|---|---|
| Infrastructure | Complete | Docker Compose with PostgreSQL, backend, and frontend |
| Database | Complete | All tables created via Alembic migrations |
| Authentication | Complete | JWT-based login with role support |
| API | Complete | All CRUD endpoints for sessions, messages, users |
| LLM Integration | Complete | Any OpenAI-compatible endpoint (default: UF Navigator), JSON mode, retry logic, structured responses |
| Dashboard UI | Complete | Session sidebar, chat, stage progress, metadata display |
| Reflection Policy | Complete | Deterministic Layer 2 FSM: five core questions with signal-based branching, loop-closure, synthesis, action-item commitment |
| Safety Monitoring | Complete | Layer 1 detects safety disclosures on every turn; a fixed reply fires and a SafetyIncident row is flagged for weekly manual dashboard review β no automated notification |
| Reliability | Complete | Both LLM calls retry once, then degrade to a plain non-alarming fallback rather than crashing; the student's raw input is committed independently so it survives a downstream failure |
| Cross-Session Memory | Complete | Each session opens on the prior session's action item and can note when the same construct recurs across recent sessions |
| Admin dashboard / export tooling | Planned | Session inspector exists; a dedicated researcher dashboard and data-export tool are not yet built |
What you can do right now:
- Log in as admin or student
- Start a chat session and have a real conversation focused on team regulation
- Watch the agent progress through the five core reflection questions, each with adaptive follow-up probing
- View Layer 1/Layer 2 classification and directive data in message metadata
- Inspect any session's full transcript and metadata on the inspect page
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β FRONTEND β
β (Next.js 14 + TypeScript) β
β βββββββββββββββ βββββββββββββββ βββββββββββββββββββββββββββββββ β
β β Login Page β β Dashboard β β AuthContext (JWT storage) β β
β βββββββββββββββ βββββββββββββββ βββββββββββββββββββββββββββββββ β
β β β
β /api/* proxy β
ββββββββββββββββββββββββββββββΌβββββββββββββββββββββββββββββββββββββββββ
β
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β BACKEND β
β (FastAPI + SQLAlchemy) β
β ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β API Routes β β
β β /auth/* β /sessions/* β /stages β /admin/* β /health β β
β ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β β
β βΌ β
β ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β FlowEngine β three-layer hybrid design β β
β β β β
β β Layer 1 (LLM) reads the student's turn, emits a signal β β
β β Layer 2 (code) decides deterministically what happens next β β
β β Layer 3 (LLM) realizes that decision as one natural turn β β
β β β β
β β prompts.py (FSM_STATES, BRANCH_SEQUENCES) βββΊ flow_engine.py β β
β β β β β
β β βΌ β β
β β llm_client.py (any OpenAI-compatible API) β β
β ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β β
β βΌ β
β ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β SQLAlchemy Models β β
β β Student β Session β Message β SafetyIncident β β
β β JSONB columns: messages.llm_metadata, sessions.evaluation_dataβ β
β ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
ββββββββββββββββββββββββββββββΌβββββββββββββββββββββββββββββββββββββββββ
β
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β PostgreSQL 15 β
β students β sessions β messages β safety_incidents β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
| Layer | Technology | Purpose | Why We Chose It |
|---|---|---|---|
| Frontend | Next.js 14 | React framework with App Router | Server-side rendering, built-in routing, great DX |
| TypeScript | Type safety | Catch errors at compile-time, better autocomplete | |
| Tailwind CSS | Utility-first styling | Rapid UI development, consistent design | |
| Axios | HTTP client | Simple API calls with interceptors for auth | |
| Backend | FastAPI | Async Python web framework | Fast, automatic API docs, modern Python async/await |
| SQLAlchemy 2.0 | ORM (Object-Relational Mapper) | Write Python objects instead of SQL queries, database-agnostic | |
| Alembic | Database migrations | Version control for database schema changes | |
| Pydantic | Request/response validation | Automatic data validation and serialization | |
| python-jose | JWT token handling | Secure stateless authentication | |
| bcrypt | Password hashing | Industry-standard password security | |
| Database | PostgreSQL 15 | Relational database | ACID compliance, JSON support, scalability |
| Infrastructure | Docker Compose | Container orchestration | One-command setup, consistent environments |
db.query(Student).filter(Student.username == "admin") vs SELECT * FROM students WHERE username = 'admin' β same query, database-agnostic (switch from PostgreSQL to MySQL without code changes).
AgenticRoboticsEvaluator/
β
βββ backend/
β βββ app/
β β βββ api/
β β β βββ deps.py # Auth and DB dependency injection
β β β βββ routes/
β β β βββ auth.py # Login, get current user
β β β βββ sessions.py # Create sessions, chat endpoint
β β β βββ stages.py # Stage registry endpoint
β β β βββ admin.py # Admin user/session management
β β β βββ health.py # Health check
β β β
β β βββ core/
β β β βββ config.py # Environment configuration
β β β βββ prompts.py # All LLM prompts and stage definitions
β β β βββ security.py # JWT and password hashing
β β β
β β βββ models/
β β β βββ student.py # User model, incl. team/subteam profile fields
β β β βββ session.py # Session with evaluation_data JSONB
β β β βββ message.py # Message with llm_metadata JSONB
β β β βββ safety_incident.py # Populated by the Layer 2 safety interrupt
β β β
β β βββ schemas/
β β β βββ auth.py
β β β βββ student.py
β β β βββ session.py
β β β βββ message.py
β β β βββ llm.py # Layer1Output / Layer3Output schemas
β β β
β β βββ services/
β β β βββ flow_engine.py # Layer 2 deterministic policy + LLM orchestration
β β β βββ llm_client.py # LLM client (any OpenAI-compatible API, default UF Navigator)
β β β
β β βββ main.py
β β
β βββ alembic/versions/ # Full migration history β run `alembic history` for the chain
β β
β βββ tests/
β βββ requirements.txt
β βββ Dockerfile
β βββ seed_admin.py
β
βββ frontend/
β βββ src/
β β βββ app/
β β β βββ layout.tsx
β β β βββ page.tsx
β β β βββ login/page.tsx
β β β βββ chat/page.tsx # Legacy chat page
β β β βββ dashboard/
β β β βββ page.tsx # Main dashboard with chat
β β β βββ [sessionId]/inspect/page.tsx # Session inspector
β β β
β β βββ components/
β β β βββ MessageCard.tsx # Chat bubble with metadata toggle
β β β βββ MetadataPanel.tsx # LLM metadata display
β β β βββ StageProgressBar.tsx # Stage progress visualization
β β β
β β βββ lib/
β β βββ api.ts
β β βββ auth-context.tsx
β β
β βββ package.json
β βββ Dockerfile
β
βββ infra/
β βββ docker-compose.yml
β βββ .env
β
βββ docs/
βββ SYSTEM.md
βββ SETUP.md
Located in backend/app/services/flow_engine.py. Each turn:
- Layer 1 (
_run_layer1) β an LLM call classifies the student's turn againstLAYER1_SYSTEM_PROMPT: pragmatic stance (cooperative / seeking clarification / active resistance / superficial compliance), signal type (thin / breakdown / success / no_event), SSRL evaluation, and a broad safety-concern flag. Retries once on failure, then falls back to an inert classification rather than crashing. - Layer 2 (
_decide) β pure Python. Checks, in priority order: a safety disclosure (highest priority, never changes the session's stage β see below), a pending safety check-in resume, synthesis/commit stage handling, loop-closure, universal interrupts (clarification, active resistance, topic drift), the Q1-Q3 Part A β Part B two-beat exchange, a coordination pivot (technical-only answers get redirected to ask about the team), then the shared branching protocol (BRANCH_SEQUENCES/BRANCH_STEPSinprompts.py) gated by a session-wide probe budget. - Layer 3 (
_run_layer3) β an LLM call realizes the chosen directive as one short turn, constrained byLAYER3_SYSTEM_PROMPT(no ground-truth verdicts, construct lock, one question mark, under 45 words, no em dashes). Retries once, then falls back to a plain non-alarming line.
Two turns are generated deterministically with no LLM call at all: the opening greeting/loop-closure and the closing goodbye.
Located in backend/app/core/prompts.py. Single source of truth for the reflection protocol:
LAYER1_SYSTEM_PROMPT/LAYER3_SYSTEM_PROMPT: the NLU classifier and NLG realizer instructionsFSM_STATES: the five core questions (Q1-Q5) plus Stage 0 (loop-closure), Stage 3 (synthesis), Stage 4 (action commit), Stage 5 (close) β Q1-Q3 each split into a guaranteed Part A/Part B two-beat exchangeBRANCH_SEQUENCES/BRANCH_STEPS: the shared thin/breakdown/success/no_event follow-up protocol applied after each core questionSESSION_PROBE_BUDGET/CUMULATIVE_PROBE_CAP: keeps a full session in a bounded turn range across a semester of weekly use- Safety interrupt copy (
SAFETY_TRUSTED_ADULT_LINE,SAFETY_CHECKIN_LINE) β fixed templates, never LLM-improvised
Located in backend/app/services/llm_client.py. Wraps any OpenAI-compatible API (default: UF Navigator) with:
- JSON mode plus a schema appended to the system prompt to structure responses
- A repair-retry loop: on invalid JSON or schema mismatch, re-prompts once with a repair instruction before giving up
LLMResultobject with token usage, response time, and attempt count
flow_engine.py layers a second, coarser retry on top of this at both Layer 1 and Layer 3 call sites β if the client still fails after its own retries, the turn degrades to a fallback response rather than propagating an error to the student.
Stage 0 Loop-closure β opens on last session's action item, if one exists
Stage 1 Orientation β explicit-criteria framing, fades across the semester
Stage 2 Core reflection β Q1 Task & Goal -> Q2 Strategy & Plan -> Q3 Monitoring & Adaptation
-> Q4 Shared Motivation & Emotion -> Q5 Evaluation
(Q1-Q3 each ask a guaranteed Part A, then Part B; each question
gets adaptive follow-up branching based on the student's signal)
Stage 3 Synthesis β mirror back what was heard, ask for one forward-planning commitment
Stage 4 Action commit β restate the ONE action item, open the floor for anything else
Stage 5 Close β deterministic, never a question
A safety interrupt can fire on any turn, at any stage, the instant Layer 1 reports a safety concern. It is not a stage β the session's current_stage never changes. Instead a fixed reply (a trusted-adult line plus a check-in question) is delivered deterministically, a SafetyIncident row is flagged for the researcher's weekly manual dashboard review, and the interrupted question is re-asked once the student is ready to continue.
1. User submits username/password to POST /auth/login
2. Backend validates credentials, returns JWT token
3. Frontend stores token in localStorage
4. All subsequent requests include Authorization: Bearer <token>
5. Backend validates token on each request via dependency injection
| Model | Table | Purpose | Status |
|---|---|---|---|
| Student | students | Users, both students and admins; also carries team/subteam profile fields | Used |
| Session | sessions | Chat session with stage tracking and evaluation_data (action item + recurring constructs) | Used |
| Message | messages | Individual messages with llm_metadata (Layer 1/2/3 audit trail) | Used |
| SafetyIncident | safety_incidents | Flagged safety disclosures, reviewed weekly by the researcher | Used |
These are the logical next steps:
-
Validation-export tool β Export the load-bearing
Layer1Outputfields (signal_type,has_social_coordination,is_team_regulation_behavior,proposed_action_item) for human-coding validation before a study begins. Field scope is decided (marked# LLM-CODEDinschemas/llm.py); the export tool itself isn't built yet. -
LLM classifier evaluation β Measure Layer 1's classification accuracy against human-coded reference labels before relying on it for study data.
# Start all services
cd infra
docker compose up
# Start with rebuild (after code changes to Dockerfile)
docker compose up --build
# Stop all services
docker compose down
# Stop and remove volumes (clears database)
docker compose down -v# All services
docker compose logs -f
# Specific service
docker compose logs -f backend
docker compose logs -f frontend
docker compose logs -f postgresdocker compose exec backend pytest -v# Connect to PostgreSQL
docker compose exec postgres psql -U evaluator -d evaluator
# Common queries
SELECT * FROM students;
SELECT * FROM sessions;
SELECT * FROM messages ORDER BY created_at DESC LIMIT 10;When the backend is running, visit:
- Swagger UI: http://localhost:8000/docs
- ReDoc: http://localhost:8000/redoc
Both frontend and backend support hot reloading:
- Backend: Changes to Python files auto-restart uvicorn
- Frontend: Next.js fast refresh on file save
| Document | Description |
|---|---|
| SYSTEM.md | Complete technical specification with data models, API contracts, and architecture decisions |
| SETUP.md | Detailed setup instructions with troubleshooting |
- Create a feature branch from
main - Make changes with clear commit messages
- Ensure tests pass:
docker compose exec backend pytest - Submit a pull request
MIT