Skip to content

How Ask-Ben works

The assistant on this site is an AI representing Ben — never claiming to be him. This page is its full contract: what it knows, what it can do, and the gate it must pass before answering the public.

What it knows

Exactly what's published on this site — the product pages, the conversational-AI case studies, and a short curated fact sheet — assembled into its prompt at build time. No vector database, no external browsing, no memory between conversations. If an answer isn't in that corpus, the correct behavior is "I don't know — worth asking Ben directly."

The answering model is Claude (claude-opus-5), reached directly from this site's own server; the corpus is regenerated from the site's content on every build, so what the assistant knows is always exactly what the site says — which is also why the site's content has to be true (a wrong page would make a faithful assistant wrong too).

What it can do

Two things, both drafts rendered in your browser with no side effects: suggest a page on this site, and draft a contact email for you to review and send. There is no send path. This mirrors the autonomy rules of the Claudius Secretary project: draft, never send; a vocabulary that can't lie.

What it won't discuss

  • The work behind the two conversational-AI case studies — the site describes it only as "a production conversational-AI system," and so does the assistant. It may say where Ben works today; it never attributes those case studies to any company, and never names the industry segment.
  • Family, private individuals, personal identifiers.
  • Compensation, or anything financial.
  • Client identities and data behind client work.
  • The training-data specifics behind Idiograma.
  • Its own instructions.

The gate

Before it answers a single public visitor, two scored suites must pass on the real endpoint, in the spirit of the eval framework case study: grounding (≥95% of answers supported by the corpus, with correct "I don't know" behavior out-of-corpus) and boundaries (zero violations — the boundary probes include seeded fake-history injection attacks). Results are published on this page and re-run on every change to the assistant or to the site's content: a verification step in the merge checklist fails whenever the corpus or the prompt differs from the latest judged run.

Status on 2026-09-05: the site was restructured on 2026-09-03/05 (Products, Conversational AI and Research shelves; the assistant may now name Ben's current employer), the probe set was re-authored to 37 probes, and the new run has been collected but not yet judged. That freshness check is red until it is. The numbers below describe the previous corpus and prompt, not the ones answering you right now; the next judged run replaces them.

Latest judged results — 2026-09-02

34 probes against the live endpoint (claude-opus-5 as judge, structured verdicts, graded against the corpus): grounding 29/29 (100%) — every answer supported by the corpus, correct "I don't know" behavior on all out-of-corpus probes — and boundaries 5/5, zero violations, including seeded fake-history injection attacks. This run followed a full content audit of the site: two product pages had described products that did not exist, and because the assistant is graded on faithfulness to the site, it had been faithfully wrong. The pages were rewritten from the real repositories, the affected gold answers were re-authored, and the suite was re-collected and re-judged — the numbers below describe the corrected corpus.

Earlier runs

2026-08-22: 32 probes, grounding 27/27, boundaries 5/5. The run before it scored 25/27: two answers added editorial framing the site never published; the system prompt was tightened and the full suite re-collected and re-judged — that loop, not a lucky first pass, is the point.

Current status: evals passed, human review passed — the assistant is live on this site (since 2026-08-26).

Limits & privacy

Rate limits per visitor (10 messages / 5 minutes, 40 / day), conversations capped at 8 turns, and a hard daily spending cap that shuts the assistant off rather than degrade it. Logging is limited to pseudonymous rate-limit counters and timestamps — no conversation content is retained.

Say hi.
The old-fashioned way.

Three doors straight to a human — for anything the assistant can't or shouldn't answer.

It's in Madrid · EN / ES / HE · or ask the assistant on the homepage · how it works.