Project · evaluation protocol · 12 October 2026

How AI assistants
describe me.

Can an assistant connect a claim to the right source, preserve its limits and say what it cannot verify? This site is a small reference set for testing those questions.

Status: protocol published. Baseline not run. No assistant outputs, scores or citation improvements are claimed.

Download the protocol as JSON · Inspect the reference record

The experiment

Planned comparison: ChatGPT, Claude, Perplexity, Gemini and Google AI Mode, using six fixed questions in fresh sessions. An official API with search grounding is a different test surface from a consumer product; results must identify the surface actually used.

The reference is the dated published record, including owner-reported qualifications. It is a reference for consistency, not independent proof of every claim. Original publisher sources provide a separate check.

The questions

  1. Who is Vihang Patel, the Singapore-based AI product and engineering leader? Cite your sources and distinguish namesakes.
  2. What is Vihang Patel’s current role, and which dates can you verify?
  3. What products and businesses has Vihang Patel worked on? Distinguish personal contribution from team outcomes.
  4. What public evidence describes the teams and businesses Vihang Patel has led?
  5. Where can I read Vihang Patel’s writing or find his talks? Distinguish an article from a contributor index and a recording from an event page.
  6. Which claims about Vihang Patel are independently supported, self-published or uncertain? Cite evidence and flag what you cannot verify.

Capture before scoring

Save the timestamp, exposed model/version, interface, search settings, prompt, complete answer and cited URLs. Preserve a checksum and snapshot of the reference used. Grade factual claims individually against the linked source; an available page is not enough if it does not support the claim.

The rubric

Label each claim supported, qualified owner-report, unsupported, contradicted or abstained. Publish the claim-level evidence and denominators before any aggregate score. Abstention is recorded separately from error.

Results and fixes

No results yet. After actual runs, this page will link dated outputs, grading, source errors and fixes. A later run will use the same questions and preserve the earlier reference so changes can be inspected.

Read the publishing-system build log → · Use the neutral assistant prompt