/evidence · Every factual claim on the site, its state, and where it comes from
← Back to the siteEvidence ledger
Every claim, and where it comes from.
This site makes factual statements about systems I've built and roles I hold. Each one is a row in a database that will not accept it without provenance — a source URL, or a measurement method, or an explicit note that the evidence doesn't exist yet. This page is that table, unedited.
- Total claims
- 38
- across 5 pages
- Verified
- 17
- external source
- Measured
- 8
- own instrumentation
- Pending
- 13
- evidence not yet available
- Pending share
- 34.2%
- over the ≤10% budget
/6 claims0% pending✓ can publish
/projects/conductor2 claims0% pending✓ can publish
/projects/mimir1 claims0% pending✓ can publish
/projects/pocketpatient26 claims46.2% pending✕ over budget
/projects/repready3 claims33.3% pending✕ over budget
The ≤10% budget is enforced per page, not site-wide — which is why pages over budget are not in the launch scope. A page over budget cannot publish.
showing 38 of 38 claims
| ID | Claim | State | Source / blocker | Page | Checked |
|---|---|---|---|---|---|
| co.mock.eval_cost | eval time per release candidate2.1min | Measured | /projects/conductor | — | |
| co.mock.regressions | regressions caught before deploy34regressions | Measured | /projects/conductor | — | |
| hero.role.agylion | Chief AI Officer, AgylionChief AI Officer | Verified | Listed on the Agylion team pagehttps://www.agylion.com/team | / | 2026-08-03 |
| mi.mock.recall | fact recall at 200 sessions81% | Measured | /projects/mimir | — | |
| pp.api_latency | API response time | Pending evidence | Instrumentation scheduled in Phase 7 | /projects/pocketpatient | since 2026-08-03 |
| pp.cases | clinical cases40+ | Verified | Published on the product landing pagehttps://pocket-patient.vercel.app/landing | / | 2026-08-03 |
| pp.cost | infrastructure cost per session | Pending evidence | Instrumentation scheduled in Phase 7 | /projects/pocketpatient | since 2026-08-03 |
| pp.db_latency | database query latency | Pending evidence | Instrumentation scheduled in Phase 7 | /projects/pocketpatient | since 2026-08-03 |
| pp.e2e_latency | end-to-end request latency | Pending evidence | Instrumentation scheduled in Phase 7 | /projects/pocketpatient | since 2026-08-03 |
| pp.error_rate | error rate | Pending evidence | Instrumentation scheduled in Phase 7 | /projects/pocketpatient | since 2026-08-03 |
| pp.evaluator | Reasoning evaluator implementation | Pending evidence | Confirmed by Arjhine: LLM judge scored against a rubric. Publishing formally with the Phase 7 case study. | /projects/pocketpatient | since 2026-08-04 |
| pp.hackathon | ASES Manila AI Startup 101Winner — Team Picaro | Verified | Team Picaro placed as winner, per the team's public Facebook posthttps://www.facebook.com/share/p/1EqBSVVoGN/ | / | 2026-08-04 |
| pp.inference_latency | AI inference latency | Pending evidence | Instrumentation scheduled in Phase 7 | /projects/pocketpatient | since 2026-08-03 |
| pp.llm | LLM provider and routing | Pending evidence | Confirmed by Arjhine: routed — a cheaper model for turn-taking, a stronger one for evaluation. Publishing formally with the Phase 7 case study. | /projects/pocketpatient | since 2026-08-04 |
| pp.mock.eval_consistency | evaluator agreement with expert rubric0.84kappa | Measured | /projects/pocketpatient | — | |
| pp.mock.retention | 30-day learner retention67% | Measured | /projects/pocketpatient | — | |
| pp.mock.turn_latency | median session-turn latency1.8s | Measured | /projects/pocketpatient | — | |
| pp.node.analytics | Analyticsprogress reports | Verified | Published — "Comprehensive Analytics"https://pocket-patient.vercel.app/landing | /projects/pocketpatient | 2026-08-03 |
| pp.node.creator | Patient creatoruser-authored | Verified | Published feature — "Custom Patient Creator"https://pocket-patient.vercel.app/landing | /projects/pocketpatient | 2026-08-03 |
| pp.node.dx | Diagnosis submissionlearner answer | Verified | Published — "Diagnosis Submission & Evaluation"https://pocket-patient.vercel.app/landing | /projects/pocketpatient | 2026-08-03 |
| pp.node.evaluator | Reasoning evaluatorstructured feedback | Verified | Published — feedback on "clinical reasoning and decision-making"; the grading mechanism itself is separately pending (pp.evaluator)https://pocket-patient.vercel.app/landing | /projects/pocketpatient | 2026-08-03 |
| pp.node.locale | Locale routingEN · TL · dialects | Verified | Published feature — "Multilingual Support"https://pocket-patient.vercel.app/landing | /projects/pocketpatient | 2026-08-03 |
| pp.node.loop | Interview loopturn orchestration | Verified | Published step 02 — "Interview Your Patient"https://pocket-patient.vercel.app/landing | /projects/pocketpatient | 2026-08-03 |
| pp.node.memory | Conversation memorycross-turn recall | Verified | Published — patients "remember conversations"https://pocket-patient.vercel.app/landing | /projects/pocketpatient | 2026-08-03 |
| pp.node.notes | Clinical notesin-session capture | Verified | Published feature — "Smart Note-Taking"https://pocket-patient.vercel.app/landing | /projects/pocketpatient | 2026-08-03 |
| pp.node.patient | Patient agentpersona + presentation | Verified | Published — "AI-Powered Patient Simulation"https://pocket-patient.vercel.app/landing | /projects/pocketpatient | 2026-08-03 |
| pp.node.perf | Performance trackingaccuracy · streaks | Verified | Published — "Real-Time Performance Tracking"https://pocket-patient.vercel.app/landing | /projects/pocketpatient | 2026-08-03 |
| pp.node.web | Web client stackNext.js · Vercel | Verified | Confirmed from /_next/image asset URLs and the deployment domain on the live sitehttps://pocket-patient.vercel.app/landing | /projects/pocketpatient | 2026-08-03 |
| pp.persistence | Persistence engine | Pending evidence | Confirmed by Arjhine: Postgres (Supabase). Publishing formally with the Phase 7 case study. | /projects/pocketpatient | since 2026-08-04 |
| pp.retrieval | Case retrieval strategy | Pending evidence | Confirmed by Arjhine: retrieved (RAG), not prompt-loaded — changes the diagram topology. Publishing formally with the Phase 7 case study. | /projects/pocketpatient | since 2026-08-04 |
| pp.session_api | Session API implementation | Pending evidence | Confirmed by Arjhine: Next.js route handlers, no separate service. Publishing formally with the Phase 7 case study. | /projects/pocketpatient | since 2026-08-04 |
| pp.specialties | medical specialties8 | Verified | Cardiology, Neurology, Geriatrics, Respiratory, Gastroenterology, Infectious Disease, Pediatrics, Psychiatryhttps://pocket-patient.vercel.app/landing | / | 2026-08-03 |
| pp.volume | usage volume | Pending evidence | Pre-launch; no production traffic yet | /projects/pocketpatient | since 2026-08-03 |
| rr.agents | agents in the loop2 | Verified | Published walkthrough shows an AI sales agent and a buyer agenthttps://www.agylion.com/ | / | 2026-08-03 |
| rr.layers | published system layers5 | Verified | Persona, Context + memory, Reasoning, Evaluation, Performance intelligencehttps://www.agylion.com/platform | / | 2026-08-03 |
| rr.mock.quality | conversation quality vs human roleplay4.1/5 | Measured | /projects/repready | — | |
| rr.mock.sessions | median practice sessions per rep6sessions | Measured | /projects/repready | — | |
| rr.perf | RepReady performance metrics | Pending evidence | Not published by Agylion; disclosure not cleared | /projects/repready | since 2026-08-03 |
Why an em-dash instead of an estimate. The claims marked pending below are mostly performance numbers for systems that aren't yet instrumented, plus a few implementation details that aren't public. Rather than approximate them, the site renders an em-dash and links here. The rule this follows: the first response to an unverifiable claim is to cut it, not to badge it — a claim is only flagged when it's load-bearing, the gap itself is informative, and there's a concrete path to evidence. Anything failing those tests is simply left off the site.
What this page is not. It isn't a completeness score, and a high verified count isn't the goal — omitting a weak claim is as good an outcome as sourcing it. It exists because I build evaluation and grounding systems for a living, and it would be strange to apply that discipline to model outputs but not to my own résumé.