Illustrated portrait of Nandini Mathan at her desk with a tablet, beside a hand-drawn AI product architecture: data, models, tools and agents, experiences, impact, with an evaluate-learn-improve loop and governance and enablement panels.
AI/ML Product Leader · Agents × Data × Domain · Bengaluru

From data science to product leadership: I build AI systems that hold up outside the demo.

I build AI products by deciding what the model does, what the system constrains, what a human inspects, and what gets measured before anyone trusts the output. Much of that muscle was built in martech: personalization, loyalty, and lifecycle marketing for brands with tens of millions of customers.

I came up shipping recommendation and pricing models to production, shaped the modeling behind a patented personalization engine, and most recently ran product for an enterprise agentic AI platform: agent orchestration, evaluation, and governed deployment. I still prototype agentic workflows and evals myself, and it keeps my product calls honest: I know when the model is the problem and when the product is.

How I think about AI products

Model, system, human: get the boundary right

Thirteen years of shipping AI, four arenas, the same discipline in each. Every card links to work that shows it.

Agentic AI platforms

Led product for enterprise agent platforms: self-service builders, agent libraries, RAG workflows, evaluation, governed deployment.

See it in practice

AI product operating models

Built the operating model that moved AI from custom delivery to repeatable product: discovery through release, metrics, feedback loops.

See it in practice

Personalization & decisioning

Led martech product lines end to end: recommendations, next-best actions, loyalty and lifecycle campaigns; AI content scaled to 8M+ personalized emails a week for QSR, airline, and fashion retail brands.

See it in practice

Responsible AI in production

Productized the controls for regulated environments: grounded outputs, human review, audit trails, PII controls, evaluation loops.

See it in practice
Experience & operating scale

Thirteen years, one through-line

From data scientist to AI product leader: the analytics-to-product bridge the whole way. Products and public client proof are named where they are public; confidential client details stay private inside the case studies attached to each role.

Jan 2026 – Present

Independent AI Product Advisor

Stealth-stage startups · Bengaluru / Remote
  • Advising early-stage teams across the AI product stack: algorithm and data modeling through product, UX, and agentic architecture.
  • Hands-on as a builder: prototyping LLM and multi-agent workflows, validating model and eval choices, shaping 0→1 roadmaps.
AI product strategyAgentic prototypingEval design0→1 roadmaps
Jun 2021 – Dec 2025

Associate Principal & Head of Product Management

ZS Associates · Bengaluru
  • Built out ZS's product management function, leveraging the firm's deep analytics and consulting muscle to productize solutions for the retail, F&B, and airlines sectors: product-market fit, a product shell to rally teams and customers around, and the discipline to make it real, then scaled it into a large cross-functional org across product, engineering, data science, design, and delivery.
  • Stayed hands-on with the technical core: prototyped agentic workflows, defined how agents were evaluated before they reached production, and applied a data scientist's judgment on data, retrieval, and failure modes to product decisions.
  • Drove product strategy, roadmap, GTM alignment, and cross-functional execution across three commercial product lines: a patented martech personalization engine for marketing and loyalty teams (Personalize.AI), an agentic AI platform (Max.AI), and a regulated-content supply chain (ZAIDYN Content); the personalization work contributed to ZS being named a Forrester Leader in Customer Analytics Services (2025).
  • Directed the shift from fragmented custom ML to a reusable agentic + GenAI platform architecture, reducing duplicated delivery effort and enabling repeatable deployment of specialized agents across regulated and commercial use cases.
  • As part of the leadership team, helped introduce the forward-deployed engineering model and shape the portfolio investment case, and partnered closely with customer success, FDEs, and clients through pilot deployments to make sure solutions succeeded in production.
Agentic AI platformsMartech & personalizationRecsys & decisioningAI governanceOrg buildingEnterprise SaaS

Products I led

Selected public proof

May 2018 – Jun 2021

Director, AI/ML Product Studio

Evolve Systems · Bengaluru
  • Founding-team role: product lead, commercial director, and solution architect at once, building an AI personalization studio from concept to production.
  • Shipped Alice (an enterprise virtual concierge for hospitality) and a retail personalization product distributed via Shopify and Magento.
  • The product, IP, and core team carried directly into ZS in 2021 to become the foundation of Personalize.AI.
0→1 productE-commerce martechRecommendation systemsConversational AISolution architecture

Selected work

Sep 2017 – Apr 2018

Data Scientist & Engagement Lead

TheMathCompany · Bengaluru
  • First-50 employee at an early-stage analytics firm, hands-on IC and internal builder simultaneously.
  • Built end-to-end as sole data scientist: designed, coded, and deployed the demand-forecasting model and pipeline for a global FMCG client: feature engineering, time-series regression, and client-facing reporting layer, no handoffs.
Demand forecastingML modelingPythonFeature engineering

Selected work

Jun 2013 – Aug 2017

Data Scientist

Mu Sigma Inc. · Bengaluru / Germany
  • Where it started: individual contributor practitioner, writing models and shipping them to production across concurrent client engagements.
  • Long-term Germany onsite with a global sportswear brand: personally designed and deployed category-affinity, purchase-propensity, and product-affinity models powering on-site recommendations, each validated in-market via A/B tests I designed and ran.
  • Sole data scientist on a dynamic pricing engine for one of the largest US secondary ticketing marketplaces: demand modeling, price sensitivity, and real-time recommendation output built and shipped end-to-end.
ML modelingAffinity / propensityDynamic pricingA/B testingSQL & Python

Selected work

Skills & stack

Technical fluency

Deep product fluency

Agentic product architecture
Planner-executor patterns, multi-agent routing, tool-use, memory & context persistence
RAG & retrieval patterns
Hybrid search, reranking, citation grounding, confidence scoring, ACL-aware retrieval
Evaluation & guardrails
Eval design, hallucination checks, groundedness & accuracy metrics, human-in-the-loop review
Responsible AI
RBAC, PII handling, audit trails, MLR/regulatory governance, HIPAA
Personalization & decisioning
Affinity, propensity, causal uplift, ranking (HinSAGE), contextual bandits
Martech & lifecycle
360 customer profiles, loyalty program mechanics, campaign orchestration, ESP/CRM integrations (Salesforce, Adobe stacks), incrementality measurement
Experimentation & measurement
A/B & incrementality testing, KPI definition, uplift measurement, event-level analysis
AI product operating models
Discovery, prioritization, release planning, success metrics, feedback loops

Hands-on / prototyping

Frameworks
LangChain, LangGraph, MCP
LLMs & GenAI
OpenAI, Anthropic/Claude, Gemini; prompt & context engineering, embeddings, multimodal, content automation
Data & code
SQL, Python, feature engineering
Retrieval infra
Vector databases, embeddings stores, hybrid search
Cloud & platform
AWS (primary), AWS & Azure Marketplace, multi-tenant SaaS, Databricks, Snowflake
Builder tooling
v0, Bolt, Replit, Cursor, Claude Code, Figma, JIRA, Amplitude
Nandini Mathan smiling and petting a black cat.
About

A product leader who started as a data scientist.

I'm Nandini, originally from Ooty, based in Bengaluru. I started as a data scientist, and it still shapes how I lead: I'm comfortable with the mechanics underneath AI products (data, retrieval, evaluation, experimentation, failure modes), but my work today is product leadership. Deciding what to build, how it should work, how it's evaluated, how teams ship it, and how customers adopt it.

I'm an ambivert: happiest deep in a problem with a small group of sharp people, and just as energized kicking off a big room at a conference. Off the clock, I trade dashboards for dirt and paws: a balcony that's slowly become a small jungle, and a cat with strong opinions who's the closest thing I have to a stakeholder review.

How I think about AI products

Three rules the builds keep proving

Models need boundaries

Decide what the model may decide

I split every workflow into what the model generates, what the system computes, and what a human inspects. Most AI product failures are boundary failures, not model failures.

Trust needs evidence

No receipt, no claim

For high-stakes or subjective output I design for traceability: citations, pins, scores, review states, visible failure modes. The verdict is computed, not narrated.

Adoption needs workflow fit

Enter the workflow, don't replace it

The best AI feature joins a workflow people already run (a chat thread, a template, a review step) instead of asking them to learn a new system.

The lab

Small builds, real product questions

I use small builds to test product questions I can't answer from demos alone. Some became working tools, some are deliberately odd. The point is the same every time: build the workflow end to end, find where the model breaks, decide what should be deterministic, and turn the lesson into product judgment. Right now that's three threads: evidence and grounding (Sift, Brand Mood Ring), automation inside constraints (the PowerPoint generator), and generative media (Blankie, Luna Whisk).

Evidence & grounding · 2026 Consumer health

Sift: a fact-checker for the health claims you hear

Can checking a health claim feel as easy as forwarding a message? Paste "seed oils cause cancer" and Sift decomposes the claim, pulls the PubMed evidence, scores study quality and directness separately, and computes a cautious verdict, on the web or straight from Telegram. The model parses and explains; it never gets to be the judge.

Read the build →
Writing

Essays on practical AI

Long-form pieces on how AI products work in production, published on LinkedIn and Medium.

Posts

Short posts

Shorter takes worth sharing.

Back to the lab
Book cover: Oh, I Can't Find My Blankie!, showing a teddy bear surrounded by an owl, bunny, cat, and dog, all searching for a lost red blanket.
Personal project · 2022–24

Oh, I Can't Find My Blankie!

Children's book Midjourney DALL·E ChatGPT Photoshop
A bedtime story about a teddy bear who loses his blankie, and the owl, bunny, cat, and dog who help him look for it. Written, argued over, and illustrated with a few friends across DALL·E and Midjourney. The whole thing was really one stubborn test: could these tools keep the same bear looking like the same bear, page after page?
The landscape

AI images went from lucky screenshots to finished work

In 2022, image models were new enough that the honest question wasn't "is this pretty" but "can it carry something finished": a dozen pages that hang together, not one lucky screenshot. Two years later that question sounds quaint: the tools reset every few months and quietly swallow whole workflows. Blankie is a snapshot from inside that shift, and the clearest way I've found to show how far it's moved.

My first experiment

A picture book is a brutal little benchmark

A children's book is a deceptively hard thing to point an image model at: a dozen-plus spreads, and the same teddy bear has to read as the same bear on every one, from every angle, in the same hand. Back then that one requirement (character consistency) was precisely what the technology was worst at, which is exactly why it was worth trying. The writing was the gentle part: plain couplets a sleepy kid could follow, like "But oh no! Where's his blankie red?", drafted with ChatGPT as a sparring partner for rhythm. Everything hard was in the pictures.

These were pure text-to-image systems with no memory between generations: each image an independent roll of the dice, with no notion of this character that carried from one prompt to the next. So the whole aesthetic had to be re-specified, in full, every single time, which made the prompt the only real lever. Ours lived as one long, exact paragraph of illustration rules (ink outline, no shadows, no hatching, a light watercolour palette, a bluish wash) where word order mattered. You generated in batches, cherry-picked the survivors, and leaned on the workarounds of the day: fixed seeds, image-to-image, inpainting, pose conditioning.

A grid of early character exploration: many different teddy bear, bunny, and owl designs generated before settling on final looks for each.
Before anyone was "the" character

A whole casting call

None of the crew, bear, bunny, owl, cat, dog, started as a fixed design. This is a slice of the rounds across both DALL·E and Midjourney it took just to find a look worth committing to, well before keeping it consistent was even the problem.

The turning point was Midjourney's character-reference feature, --cref. For the first time you could hand the model an approved portrait and say, in effect, this one, tuning how hard it listened with --cw. It was the single most useful thing in the whole project. We settled on a Teddy everyone agreed on, then pointed every subsequent scene back at that reference.

The approved reference portrait of Teddy: a small brown bear in an orange zip-up jacket and denim shorts, mid-jump, in hand-drawn ink and watercolour style.
The one we pointed everything back at

Teddy, locked

This portrait became the --cref for the rest of the book. The jacket, the proportions, the expression, all of it had to survive into every other page, at every other angle, for months.

The prompt that made him "Teddy"

A highly simplified, hand-drawn ink illustration of a small brown teddy bear named Teddy, wearing an orange zip-up jacket and denim shorts. The style is minimalistic, with minimal details and no shadows, no hatching, outlined in black ink for a crisp, clean appearance. The colors used are very light, watercolour, and pastel, emphasizing a gentle aesthetic. Forced perspective, Teddy mid-jump, arms out, looking straight at camera, a few motion marks trailing his feet. --v 6.0

From there, every scene reused that same style paragraph with --cref aimed at the portrait. It helped. And it still wasn't enough.

Four Midjourney generations from the exact same prompt and character reference, all showing Teddy on a ferris wheel, each rendered in a noticeably different style.
Same prompt, same reference

Still four different bears

This is the honest limit of that era, in one batch: identical prompt, identical reference, identical weight. One reads almost photoreal, one is a loose painterly sketch, two are close but not twins. The reference narrowed the range; it never closed it. Every batch still ended with a human picking the one that belonged in the book.

Even the keepers weren't done. Each needed real Photoshop work: a stray extra finger, lighting matched across spreads, text composited into the art, print files prepped. The "AI-generated book" was, in practice, a lot of manual craft stitched around the generations. And the sharpest editor in the whole process wasn't a model at all: Priya went through the draft page by page, flagged that naming every character right before bed was one too many for a drowsy toddler so that page got cut, then read the whole thing aloud to her own kid and reported back on what landed.

WhatsApp message from Priya giving page-by-page feedback on the book draft. WhatsApp message from Priya continuing her feedback, followed by a reply saying she'd read it to her kid the next day.
Three finished spreads from the book, showing Teddy getting ready for bed, realizing his blankie is missing, and looking for it around the room.
What shipped

The finished pages

All that casting and re-generating, in service of three quiet spreads: Teddy ready for bed, Teddy realizing blankie's gone, Teddy looking here, then there. It went out as a real, self-published, if-you-lose-it-you-owe-me-a-copy book, with a calming narrated version following on YouTube much later.

The current wave

The gap that closed in three years

Almost every hard-won trick in the book has since been absorbed into the models themselves. The re-typed style paragraph, the --cref reference, the seed-locking, the batch-and-cherry-pick, the inpainting, the Photoshop salvage. That entire apparatus existed to paper over one gap: the model couldn't hold a character in its head.

Now it largely can. Multimodal, in-context generation lets you hand a model your character and your instructions in the same breath, keep it consistent across a whole set without a special flag, and, most tellingly, edit in place by conversation: "same bear, now on a swing, keep the jacket," instead of rolling the dice again. That moves where the skill lives: when consistency is nearly free, taste and art direction matter more than prompt-craft. Re-made today, Blankie would be closer to an afternoon than the months it took.

Why it stays with me

Made for bedtimes, not a portfolio

This was never a portfolio exercise. It got made to be read to actual kids at actual bedtimes, argued over on WhatsApp threads, and revised because a real toddler lost interest on the wrong page. It's easy to track this shift from release notes and demos; Blankie is my argument for the opposite: you only really feel how far the tools have moved by trying to make something finished with them, then trying again a year later.

The book is finished. The field it came out of isn't, and that's the part I can't stop watching.

Back to the lab
Personal project · 2023

Inter-Dimensional Adventures of Luna Whisk

GEN:48 finalist RunwayML DALL·E Suno AI ElevenLabs

A short film made with a few friends in 48 hours flat, for the first-ever edition of RunwayML's GEN:48 challenge: one location, one character, and one event handed to every team at the same moment, with a hard deadline and no extensions. Ours made the finalist list.

Watch it on RunwayML
The pace

AI video moves faster than almost anything I've worked on

Text-to-video has gone from a party trick to something you can almost put in front of an audience, in about two years. Whole categories of work (shorts, ads, explainers, music videos) now get roughed out by people who'd never go near a render farm, and the state of the art turns over on a scale of months. I've been making things inside that churn since close to the beginning; the first one that felt real was a 48-hour film called Luna Whisk.

The 48-hour build

Luna Whisk, start to finish in a weekend

GEN:48 was RunwayML's first 48-hour film challenge. Every team gets the same three ingredients at the same moment (a location, a character, an event), then races a hard deadline with no extensions. Ours became an interdimensional cat named Luna Whisk, made with a few friends over one weekend.

The pipeline was whatever we could stitch together. DALL·E generated the interdimensional worlds. Runway's image generator handled character shots, fed reference images to keep Luna on-model. Its Gen-Turbo model turned those stills into motion. An original score came from Suno, sound effects from the Epidemic Sound library the contest handed participants, and narration from an AI-cloned voice in ElevenLabs. A friend who does this for a living wired the tools together, knew which warped takes to throw out, and did the final color and edit, exporting minutes before the cutoff.

This edition predated most of today's tooling, so those five tools were, between them, close to the entire stack. And the hard part was never the look. It was continuity. Nothing enforced that the cat in one shot was the cat in the next, so "the same cat" meant close enough, picked by eye.

The pipeline, drawn out

Six tools in, one film out

Every box below was a separate product, and every arrow was a hand-off we made by hand. That topology is why continuity was the hard part: nothing downstream knew what the step before it had produced.

2023 · FIVE AI TOOLS, WIRED BY HAND IMAGEIMAGEMOTIONSCOREVOICEEDIT DALL·ERunwayRunwaySunoElevenLabsmanual worldscharacter stillsstills → motionsoundtracknarrationcut · grade · export continuity held by hand · reference + pick-by-eye — dashed hand-off = a file exported from one tool, re-imported into the next, by hand NOW · COLLAPSING TOWARD ONE MODEL GENERATEFINISHOUT Unified video modelHuman editFilm prompt + references · native audio · one passlighter cutfewer tools one shared state · not six exports
A filmstrip of early Runway attempts at the cat character: a reference photo, two illustrated results in a hooded robe, a photorealistic kitten on a lily pad, and a painterly kitten in a teacup, each in a noticeably different style.
Before we found her

What "a cat character" got back

A few of the early passes: a shutterstock-style mascot in a hooded robe, a photoreal ginger stray on a lily pad, a painterly kitten curled in a teacup. None of them were the white cat with amber eyes and a glowing third-eye gem that eventually shipped. Landing on a reference worth locking in took more rounds than getting a good scene once we had one.

Luna Whisk reaching toward a glowing gem on a lily pad in a mossy, sunlit interdimensional forest.
Keeping her on-model

Luna Whisk, mid-scene

One of the reference-anchored character shots from Runway: the same white cat with amber eyes and the glowing third-eye gem, dropped into a new backdrop and still recognizably herself.

Luna Whisk leaping through a glowing dimensional rift above a lit floor grid, with a striped tail and softer markings than in the earlier character shots.
The consistency problem, in one frame

Same character, different pass

This one came out of a different generation pass, and it shows: a striped tail and softer markings, next to the clean white-and-orange coat in the earlier stills. Nothing in the pipeline enforced one source-of-truth reference across every scene, so "the same cat" meant close enough, picked by eye, more often than it meant identical.

A few weeks later

Then, the finalist email

Email from Runway Events confirming the film's selection as a GEN:48 finalist.
Under the hood

From stitched clips to something closer to a world model

The reason all that stitching was necessary is worth naming. Runway in 2023 was really an image model taught to move: each shot was denoised on its own, with no memory of the one before it, so it had no concept of the same cat, only a plausible next frame. Almost everything interesting since has come from dropping that one assumption.

The current wave of models treats a clip as a single object in space and time instead of guessing frame by frame, which is most of the gap between the old "AI melt" and a shot that holds a coherent character and light source across ten seconds. References are a first-class input now rather than a folder of screenshots; models that score their own audio fold the Suno-and-ElevenLabs step back into the render; the five-tool pipeline is collapsing toward one. It's nowhere near solved (long-range consistency and directing a shot are still open), but "keep this one cat consistent" went from the whole challenge to almost a checkbox in two years.

The same shift, as a scorecard

What moved, dimension by dimension

Dimension2023 · what I worked withThe shift since
ArchitectureImage diffusion, one clip at a timeA spacetime diffusion transformer: the clip modelled as one object, not guessed frame by frame
Temporal unit2–4s clips, each generated alone10s+ coherent shots in a single pass
CharacterReference images + pick-by-eyeReferences as a first-class, built-in input
Cross-shot continuityNone enforced; held by a humanEmerging shot & scene memory
AudioSuno + ElevenLabs, stitched in afterNative audio & dialogue in the same render
Complex motionWarps and morphs, "AI melt"Rough physics & object permanence
PipelineFive tools + manual glueCollapsing toward one or two

Reference points for the right column: Sora (spacetime patches, 2024), Runway Gen-3/Gen-4, Google Veo, Kling, Luma Dream Machine.

Why I keep pulling on it

Where a capability actually is, versus the demo

Here's the professional reason I build things end to end like this: it's how I learn where a capability is: not where the demo puts it, but where it breaks, what's still held together by people, and which of those seams the next model quietly absorbs. Reading the real frontier against the marketed one is most of what I bring to AI product work, and it doesn't come from a changelog.

The un-strategic reason is simpler: I find this genuinely amazing, and I'd rather stay close to the thing than only read about it. Luna Whisk is what came out of following that: a weekend with friends, made to see what would happen.

Luna promises endless interdimensional cuddles in return. 🐱

Back to the lab
Personal project · 2025

A PowerPoint generator, built the year before the models could

Agentic AI Python python-pptx OOXML GPT-4o

In early 2025, "make me a deck about X" was still an open gap. You could get an LLM to write slide copy, but what came back looked like AI slop. Centered text on a white background, none of the craft of a deck someone designed. So over a couple of weekends I built the opposite: not a model that designs slides, but a small crew of agents that fill ones a designer already got right. It worked end to end, rendering real decks with real content and visuals.

a weekend rabbit-hole that shipped working decks, then the labs made it a button and I happily let them.

The bets, up front

Product questionCan agents generate useful decks inside an existing design system?
UserA team that needs a decent first draft without wrecking the template.
Design betDon't ask the model to design the slide. Give it slots, budgets, and constraints.
System betExtract the .pptx into a layout contract, generate into that contract, reject anything that overflows.
What brokeThe model treated character limits as vibes until the fit check made them measurable.
What I'd changeKeep the contract, swap the writer. The loop, not the prompt, is the part that aged well.

A good slide is two things wearing one coat: a design (where the title sits, what the palette is, how many characters a box holds before it looks cramped) and the content that lands inside it. Models in early 2025 were strong at the second and hopeless at the first.

So I stopped asking them to do the first at all. I bought a professionally-designed template (a light modernist pitch deck, thirteen slides) and treated the design as fixed. The whole system had one job: read that deck into a spec precise enough that agents could pour new content into it without touching a single design decision. The template is the contract; the agents just honor it.

Frey Munch
Product
Pitch
deck
SUBTITLE · budget ≤324 chars
CENTER_TITLE · budget ≤1314 chars
Slide 1 of the real template, rebuilt in its own theme (mint #3AEFCC, Arial Black titles). The dashed boxes are the two text slots the extractor found, each tagged with the exact character budget, pulled from the deck, that the agents have to write within.
The pivot that mattered

First I asked the model to read the deck. Then I stopped guessing.

The obvious first build was to dump every slide's shapes into GPT-4o and ask it to describe the template as JSON. It demoed fine and failed in practice: the model hallucinated placeholders that weren't in the deck and dropped ones that were. It was guessing at a file it couldn't see.

script.ipynbPython
# v1 (rejected) — ask the model to infer the template from a shape dump
resp = openai.chat.completions.create(
    model="gpt-4o",
    response_format={"type": "json_object"},
    messages=[
        {"role": "system", "content": "You are an expert in slide design."},
        {"role": "user",   "content": prompt},   # raw shapes + instructions
    ],
    temperature=0.7,
)
# Failure: invents placeholders not in the deck, ignores ones that are.

So I threw it out. A .pptx is just a zip of XML, so v2 reads the actual document instead of describing it: every text frame's position, font, size, weight, color, alignment, and bullet levels, pulled from the drawing-markup, with python-pptx as a fallback when the XML stays silent. No guessing; the spec now describes the deck that exists.

utils.py · _get_font_infoPython
# v2 — read the drawing-markup directly (source of truth, not a guess)
NS = {"a": "http://schemas.openxmlformats.org/drawingml/2006/main"}

for p in txBody.findall(".//a:p", NS):
    defRPr = p.find(".//a:pPr/a:defRPr", NS)
    if defRPr is None: continue
    if "sz" in defRPr.attrib:                 # size stored x100
        font_size = float(defRPr.attrib["sz"]) / 100.0
    latin = defRPr.find(".//a:latin", NS)
    if latin is not None:
        font_name = latin.attrib.get("typeface")
    srgb = defRPr.find(".//a:solidFill/a:srgbClr", NS)
    if srgb is not None:
        text_color = "#" + srgb.attrib["val"]

Everything the extractor produces lands in one typed record per text frame. Get the schema right and the reading half and the writing half can be built and swapped independently. They only ever agree on this shape, and max_chars is the line that makes the contract enforceable:

utils.py · dataclassesPython
@dataclass
class TextFrame:
    id: str                       # unique per frame (uuid4[:8])
    placeholder_type: str = None  # CENTER_TITLE, SUBTITLE, BODY, ...
    text: str = ""
    left: float = 0.0; top: float = 0.0; width: float = 0.0; height: float = 0.0
    max_chars: int = 0            # the constraint the agents must respect
    font_name: str = None; font_size: float = 0.0; font_bold: bool = False
    text_color: str = None; horizontal_alignment: str = None
    bullet_points: list = field(default_factory=list)
    level: int = 0               # 0 title, 1 subtitle, 2 body

Built from six pieces, each doing one job

python-pptx

The friendly path into OOXML: iterate slides, shapes, placeholders; read geometry in EMU, convert to inches, and at the end write filled copy back in place.

xml.etree.ElementTree

The escape hatch. When the high-level API returns None for inherited styling, I drop to the raw a:defRPr markup and read it myself.

dataclasses

Typed TextFrame / SlideTemplate as the spec schema; asdict() serializes straight to JSON. The schema is the interface between extractor and agents.

collections.Counter

Real decks are messy: runs disagree on font. Counter resolves the dominant font, size, and color per box.

openai · gpt-4o

The engine behind the planner and content agents, in json_object mode, alongside the rejected v1 kept in the repo as a documented dead-end.

uuid

Short ids per frame so an agent can address one specific slot (fill frame b5ee39cd) without positional ambiguity.

flowchart LR
  A["Designed .pptx<br/>(bought template)"] --> B["PPTTemplateExtractor<br/>python-pptx + raw OOXML"]
  A --> C["ThemeExtractor<br/>colors / fonts / master styles"]
  B --> D["template_data.json<br/>per-slide spec + char budgets"]
  C --> E["theme_data.json"]
  D --> F{{"The contract"}}
  E --> F
  F --> G["Agent loop<br/>plan / write / fit-check"]
  G --> H["Rendered .pptx<br/>design untouched"]
Deterministic where the file has the answer; agentic only where it genuinely doesn't.
The two hard parts

Inheritance, and giving every box a budget

Most slides don't state their own formatting. A title box inherits its font from the layout, which inherits from the master, which falls back to the theme. Read the shape alone and you get None for nearly everything, so the extractor walks the same cascade PowerPoint resolves at render time, stopping at the first level that answers.

flowchart TD
  S["Shape run properties"] -->|null?| L["Layout placeholder"]
  L -->|null?| LD["Layout default font"]
  LD -->|null?| M["Master default font"]
  M -->|null?| T["Theme font scheme"]
  T -->|null?| F["Hardcoded fallback<br/>Arial 12pt, #000"]
  S -.first hit.-> R{{"Effective formatting"}}
  L -.-> R
  LD -.-> R
  M -.-> R
  T -.-> R
  F -.-> R
A box that "says nothing" still resolves to the font it renders in.

The second part is the one that makes the whole thing enforceable. For each frame the tool estimates how much copy fits from the box's real size and font, and writes that number into the spec as a hard limit. Downstream, the content agent is told "this subtitle gets ~324 characters," not "write a subtitle." That budget is what later lets a fit-check reject a slide that overflows.

utils.py · _estimate_max_charsPython
def _estimate_max_chars(self, width, height, font_size):
    # avg glyph advance ~= 60% of point size (heuristic, not true metrics)
    chars_per_inch = 72 / (font_size * 0.6)
    chars_per_line = max(int(width * chars_per_inch), 10)
    line_height    = font_size * 1.2 / 72        # points to inches
    max_lines      = max(int(height / line_height), 1)
    return chars_per_line * max_lines           # the box's copy budget
How it filled & rendered

A small crew of agents, one feedback loop

With the contract in hand, generation stopped being a single prompt and became a loop of narrow agents, each with one responsibility. Give it a topic and the spec, and:

A planner chooses which template slides to use and orders them into a story. A content agent writes copy into each slot, one frame at a time, working against that frame's max_chars. A visual agent generates the charts and imagery, reading from the same theme palette the extractor pulled so nothing clashes. Then a fit-check measures every box against its budget, and anything over goes straight back to the content agent for a tighter rewrite. Only once every slide passes does the render agent write the filled spec back into a real .pptx through python-pptx, editing runs in place so the design survives untouched.

flowchart TD
  U["Topic + template spec"] --> P["Planner agent<br/>choose slides, order the story"]
  P --> C["Content agent<br/>write copy into each slot"]
  C --> V["Visual agent<br/>chart / image in theme palette"]
  V --> K{"Fit check<br/>every box within max_chars?"}
  K -->|over budget| C
  K -->|passes| Rn["Render agent<br/>python-pptx writes .pptx"]
  Rn --> D["Finished deck<br/>design untouched"]
The fit-check is the whole trick: the loop, not the prompt, is what forces the copy to respect the design.
agent output
The opportunity
Enterprise AI is past the demo
  • Buyers now ask for governance, not novelty.
  • Adoption, not accuracy, is the bottleneck.
  • The winners ship boring, reliable systems.
A content slide the crew planned, wrote, and rendered. Each line trimmed to its box's budget by the fit-check, with a simple chart drawn from the theme's own accent palette.
Challenges faced

What fought back

Getting a model to respect a hard limit. An LLM treats "≤324 characters" as a polite suggestion. No amount of prompt-tuning made it reliable. The fix was structural, not verbal: generate, measure, and bounce anything over budget back for a rewrite. The fit-check loop is what turned a suggestion into a constraint, and it's the piece that made everything else trustworthy.

Rendering back without wrecking the design. Writing generated copy into a slide is deceptively easy to get wrong: replace a text frame wholesale and you lose the exact fonts, colors, and spacing the template existed to provide. The render agent had to edit runs in place, changing the words while leaving every formatting attribute the extractor had so carefully resolved.

Visuals that looked like they belonged. A perfectly-written slide still looks stitched together if the chart shows up in stock blue. Feeding the extracted accent palette into the visual agent kept generated graphics on-theme: small thing, big difference in whether a deck reads as designed or assembled.

The budget is still a heuristic. The character estimate approximates glyph width; it doesn't know an i from a W, so it's occasionally off. I shipped it anyway, because the fit-check downstream caught the misses, an imperfect budget was good enough to build a reliable loop on. Real per-typeface metrics would tighten it, and that's the first thing I'd redo.

What changed

The models shipped it, and got the framing right

I got it working end to end: a topic in, and the agents planned the deck, wrote to every box's budget, generated visuals in the theme palette, and rendered a real .pptx with the design intact. A few months after I had it producing decks, Claude and Gemini shipped native deck and document generation, and any reason to keep building a bespoke pipeline evaporated. That's the right outcome: if a frontier model does the whole job well, a weekend project shouldn't compete.

But it was satisfying to watch, because the thing they got right is the thing this was built around: quality comes from working within a strong design system, not from asking a model to conjure one each time. What stays with me isn't the code, it's the framing: find the taste-dependent half of a problem, make it a fixed input, and let agents do only the part they're reliable at, inside a loop that checks their work. That instinct outlives the prototype, which is the whole point of building small things at the edge of what's possible. You learn where the edge is with your hands on the file, not from the release notes.

Don't ask a model to design the slide. Ask a crew of agents to fill the one a designer already got right, and check their work.

Back to the lab
Personal project · Multimodal AI · Creative critique

Brand Mood Ring

Upload a creative. Find out what it really says.

Multimodal AI Vision LLM Agentic Built & running

Brand Mood Ring is a playful interface around a serious product question: can a multimodal system critique a marketing asset in a way that is grounded, useful, and emotionally legible? Drop in a poster, ad, landing-page screenshot, or a raw HTML email, and it returns a creative diagnosis: what it feels like, who it’s really for, and what it’s accidentally saying. The fun is in the surface. The discipline is underneath: every read has to point back to a quoted line, a visual region, or a palette value.

less “brand compliance score,” more “your asset is wearing the brand colors, but emotionally it has wandered into webinar territory.” started as one curious weekend; now it’s a working app, and every screenshot on this page is a real scan. the person I picture using it: someone about to ship a campaign who half-suspects their asset is off but can’t name why.

The bets, up front

Product questionCan a multimodal system critique a marketing asset in a way that is grounded, useful, and emotionally legible?
UserSomeone about to ship a campaign who half-suspects the asset is off but can't name why.
Design betA playful surface (creatures, roasts, a mood ring) so the critique lands as a read, not a report card.
System betA Python scout computes the receipts (OCR regions, palette values) before any model speaks; a referee bounces claims without evidence.
What brokeThe critic sounded equally confident whether or not it was grounded. The referee pass exists because of that.
What I'd changeAdd first-glance saliency signals so the 3-second read is measured, not asserted.
Why I made this

I started with a dumb question

I wanted to know if a vision model could read vibe: not “this is an email with a button,” but “this is trying to be your funny friend, and mostly pulling it off.” One weekend turned into several, because the diagnoses kept coming back sharper (and funnier) than I expected.

What kept me building: the gap read turned out useful. Most creative isn’t bad, it’s mismatched. It intends one thing and signals another, and naming that gap in plain, slightly rude language turns out to be something a multimodal model is weirdly good at. Somewhere along the way the prompt chain grew a scout, a referee, and opinions about evidence, and became an actual app.

What I built

So I built a scanner with four reads

Vibe read
It reads the asset’s visual style, emotional tone, implied audience, and first 3-second impression.
Signal gap
It compares what the creative seems to intend with what it actually signals.
Cliche detector
It calls out the category’s overused moves: AI gradients, vague transformation claims, template layouts, weak CTAs, generic proof.
Fix directions
And it suggests fixes: sharper copy, clearer hierarchy, stronger trust signals, alternate creative directions.
Scan one

I fed it a Zomato email, Goblin Mode on

This is the app mid-diagnosis: Streamlit on top, a crew of Claude agents underneath. Left pane: the uploaded email with numbered evidence pins, drawn from OCR boxes a Python scout measured before any model said a word. Top right: the crew tracker, tagger to critic to referee to goblin, ticking green as each agent lands. The verdict: the Contraband Cookie Bandit.

The Brand Mood Ring app after scanning a Zomato promo email: the uploaded email on the left with numbered evidence pins, the crew progress tracker and the Goblin Mode roast on the right.
Every claim points at a pin, and every pin is a real pixel region the scout measured before the model spoke.

That one scan fills six tabs. Each is a different read on the same pixels, and every screenshot below is real output, from this Zomato scan or the Olaplex one further down:

Signals tab: trust receipt, copy versus visual, and the referee verdict.
Signals
How it earns trust, whether copy and design agree, and the referee’s grounded-or-not verdict.
Tags tab: four lenses of chips for content, visual, emotional, and audience.
Tags
The raw four-lens vocabulary, content, visual, emotional, audience, that every later read builds from.
Audience tab: persona reactions to the Zomato email.
Audience
Four rooms role-played: who lights up, who winces, and who it repels on purpose.
Remix tab: a more-wholesome rewrite of the Zomato cheat sheet.
Remix
Pick a dial and it rewrites in-voice; the critic re-scores until the brand DNA survives.
Share card tab: the shareable diagnosis card for the Contraband Cookie Bandit.
Share card
The whole diagnosis compressed into one shareable creature, with the full JSON a click away.
Goblin Mode card: a sharper roast paired with a practical translation.
Goblin Mode
A sharper, mildly rude roast, always paired with a translation back into a to-do list.
The Bandit’s read: copy and visual a full match, the referee ruled it grounded, and its one flaw is the green button, “a utility bolt on a hand-drawn machine.”
Scan two

Then I pointed it at quiet luxury

To prove the ring reads vibe and not a checklist, I picked the Bandit’s exact opposite: an Olaplex nurture email. Premium, clinical, five thousand pixels of bond science. The ring named it the White-Coat Rapunzel, and the flaw it found is structural, not tonal: one email wired for two customer journeys, a welcome gift up top and a cart-recovery button right under the science. This is also the scan where the referee earned its keep, bouncing the critic’s first draft for ungrounded claims and forcing a redo.

The app after scanning an Olaplex nurture email: pinned evidence on the email, the crew tracker, the Goblin Mode roast, and the creature verdict, the White-Coat Rapunzel.
Same ring, opposite creature. Pins 1 through 5 mark the receipts: the wordmark, the molecular claim, the FINISH SHOPPING button, the 15% welcome code.
The referee card for the Olaplex scan showing the had-ungrounded-claims-redone badge with a claim-by-claim audit.
“Had ungrounded claims, redone.” Most honesty layers only live in the architecture diagram. This one bounced a real claim and forced a redo.
Under the hood

How I built it: receipts first, opinions second

Ingest normalizes the input: images pass straight through; raw HTML gets rendered to pixels by headless Chrome and parsed for its DOM facts, so the model reads what the email says and what it looks like. Then a deterministic Python scout runs OCR and k-means clustering to produce the receipts: text blocks with ids and pixel boxes, a palette with exact percentages. Only after that does the crew speak: a tagger sorts, a critic interprets, a referee bounces any claim the scout can’t back up, and the goblin and remixer go last.

flowchart LR
  U["Upload<br/>image or HTML"] --> IN["Ingest<br/>HTML rendered + DOM parsed"]
  IN --> SC["Scout, plain Python<br/>OCR receipts + palette"]
  SC --> TG["Tagger agent<br/>four-lens tags"]
  TG --> CR["Critic agent<br/>vibe, gap, cliches"]
  CR --> RF{"Referee<br/>every claim has a receipt?"}
  RF -->|grounded| OUT["Structured JSON<br/>renders the six tabs"]
  RF -->|bounced, redo once| CR
  CR -.-> GB["Goblin agent<br/>roast + translation"]
  CR -.-> RX["Remix agent<br/>dial-driven rewrites"]
  GB -.-> OUT
  RX -.-> OUT
The referee is the honesty layer, and it isn’t decorative: on the Olaplex scan it bounced the first read, fed the critic its objections, and got a grounded redo back.
The scout's palette receipt for the Zomato email: five hex swatches with exact percentages, led by #1e1e21 at 36.5 percent.
The scout is deterministic: same pixels, same receipts. “#1e1e21 at 36.5%” is a measurement, not a vibe, which is what lets the referee audit the vibes.
The scan sidebar: file uploader, sample email button, Goblin Mode toggle, scan button, and a crew backend readout showing claude code.
Two interchangeable backends: the Anthropic API with schema-enforced structured outputs, or the Claude Code CLI running on a plain subscription. The sidebar tells you which crew showed up.
flowchart LR
  D["Dial picked<br/>e.g. more wholesome"] --> RX["Remix agent<br/>drafts the rewrite"]
  RX --> CR["Critic agent<br/>re-scores the draft"]
  CR -->|voice drifted| RX
  CR -->|same DNA, new angle| SHIP["Rewrite ships"]
Remix is a loop, not a one-shot: the critic re-scores every draft, two rounds max, until the brand voice survives the rewrite.
The rules that keep it honest: the crew keeps observation separate from interpretation, never fakes precision (“make it pop” doesn’t appear in its vocabulary), critiques the asset and never the person who made it, and won’t invent a brand-fit verdict it has no basis for. Above all, every claim has to trace to a receipt: a quoted line, a pinned region, a palette value. No receipt, no claim. That last rule isn’t a prompt suggestion, it’s the referee’s job description.
one scan, trimmedJSON
{
  "creature": { "name": "the Contraband Cookie Bandit",
                "chips": ["conspiratorial", "meme-fluent"] },
  "scores": { "clarity": "High", "trust": "Medium", "distinctiveness": "High" },
  "palette": [{ "color": "#1e1e21", "percentage": 36.5 }, "..."],
  "evidence_pins": [{ "pin": 1, "receipt_id": "R1" }, "..."],
  "referee": { "grounded": true }
}
The eval problem

How I test something with no answer key

Taste has no ground truth, so I test around it. The bar is three properties I can measure:

Stable
Ten scans of the same asset should return the same creature, dials within a notch. A diagnosis that changes on refresh is a horoscope.
Agreed
A blind panel of humans reads the asset, the ring reads it too, and the gap calls get compared. Wording can differ; the diagnosis shouldn’t.
Grounded
Every claim traces to a receipt. The referee enforces this at runtime instead of trusting the prompt to behave.
Grounded already ships: the Zomato read passed clean, the Olaplex read got bounced and redone. Stable and Agreed are the open evals.
Runs today
Everything on this page: ingest (images and raw HTML), the scout’s receipts, the tagger, critic and referee crew, Goblin Mode, remix dials, the share card, full-scan JSON download, two model backends. Every screenshot is the app, mid-scan.
Still a drawing
A public deployment you can poke without cloning the repo, the stability and blind-panel evals run at scale, and everything in the roadmap below.
The roadmap

What I’m building next

Plug it into neural signals the big one
Right now the ring guesses what people notice in the first three seconds. Real attention data could grade that guess. Eye-tracking, or the predictive-attention models trained on it, would turn the First 3s read from taste into something you can test. It also raises the eval bar: Agreed stops meaning a panel of humans and starts meaning actual attention.
eye-tracking heatmaps predictive attention models saliency benchmarks
A public creature gallery
Every scan, with permission, joins a browsable zoo of diagnosed creatives. The shareable card becomes the growth loop.
Scan in place
A browser extension that reads any live page where it stands. No screenshots, no uploads.
Let it argue with itself
Two personas read the same asset and debate before the ring commits. Disagreement becomes a confidence signal.
Same toy. Real instruments behind the guesses.

It doesn’t ask whether your creative is good. It asks what your creative is making people believe.

python + a claude agent crew · runs on my machine · public deploy on the roadmap ← Back to the lab
Back to the lab
Personal project · Evidence AI · 2026

Sift

Fact-check the health claims you hear.

Evidence AI FastAPI OpenClaw Next.js Telegram PubMed

A friend forwards a reel ("seed oils cause cancer") and you want to know if it’s true. The honest answer is buried in a research database built for scientists, not for someone standing in a kitchen. Sift is the layer I built in between. You paste the claim you heard, and it breaks the claim apart, pulls the evidence, weighs how good that evidence actually is, and hands back a calm, source-backed verdict, on the web or straight from Telegram. It runs end to end, and every screenshot on this page is the real thing.

the whole flow works, web and telegram. the store ships seeded around 25 of the claims people actually repeat; when something new comes in and the local shelf is thin, it goes out to pubmed live and keeps what it finds. the person I picture using it: someone who just got sent a scary health clip and wants one honest answer before dinner.

The bets, up front

Product questionCan checking a health claim feel as easy as forwarding a message?
UserSomeone who just got sent a scary nutrition reel and wants one calm answer before acting on it.
Design betStart from the claim, not the paper. Keep the front door to one input.
System betThe model parses and explains; the verdict is computed from structured, scored evidence.
What brokeStudy quality and claim directness looked like one axis until the scoring forced them apart.
What I'd changeGrow the evidence graph past the seeded claims. The ontology, not the chat, is where this compounds.
Why I made this

It started with a claim I couldn’t check fast enough

Someone sent me "creatine causes hair loss" as a fifteen-second clip. Checking it properly means knowing which of forty thousand PubMed hits actually speak to the question, which ran on people instead of cells, and which are strong enough to change what you’d do about it. That’s a researcher’s afternoon, and nobody watching the clip has one to spare.

Consensus and tools like it already proved people will use an evidence search engine as long as it talks plainly. I wanted something narrower and more opinionated: not a search box over papers, but a claim checker. You hand it the exact sentence you heard, and it does the translation, because people think in claims and PubMed thinks in papers. Everything Sift does lives in the gap between those two, and it always starts from the claim, never from the paper.

The front door

One box, and whatever you just heard

I kept the home page deliberately plain: one input, a few example claims so you’re not staring at an empty box, and a single button. Underneath it I put a wall of trending checks with their verdict chips already attached, so the vocabulary teaches itself before you type a word. You see "Unsupported" and "Likely supported" and "Overstated" sitting next to real claims and get the idea. There are no accounts and no dashboard. All the complexity I could push backward, I did; the page you land on doesn’t show any of it.

localhost:3000 · the running build
Sift home page: a single claim input with example chips, trending checks with verdict chips like Unsupported and Likely supported, and a how-it-works explainer.
The real home page. I even put the grounding rule in the "how it works" box at the bottom, in plain sight: the model never decides the verdict.
On Telegram

Catching the claim where you actually heard it

Health claims almost never arrive while you’re sitting at a browser. They show up mid-conversation, in the chat you already have open, so I put Sift there too. Send it a single claim and it replies with the verdict, the confidence, a couple of plain sentences, and a link to the full report if you want to go deeper. Paste a whole influencer rant and it does something I like more: it reads the blur, pulls out the separate claims hiding inside it, and asks which one you want checked. Both of these screenshots are the actual bot, on my phone.

The Sift Telegram bot answering a single claim: the user asks whether vegetable oils cause cancer, and the bot replies Unsupported as a broad claim, with moderate confidence, a plain-English explanation, a link to the full report, and a not-medical-advice line.
One claim in, one calm answer back: the verdict, the confidence, why, a link out, and the disclaimer. No scarier than it needs to be.
The Sift Telegram bot handling a bundled message that mixes three claims about seed oils, pressure cooking, and paneer protein; the bot replies that it found three checkable claims, lists them, and asks the user to reply 1, 2, 3, or all, then begins checking.
A bundled message about seed oils, pressure cooking, and paneer gets split into three separate claims to check, instead of answering the whole thing as one blur.

The bot isn’t a one-off integration I hand-wired. It runs on OpenClaw, a local-first, messaging-first agent runtime, which means the Telegram gateway, the conversational state that remembers which claim you picked, and the agent loop all come for free and all stay on hardware I control, which matters more than usual for a health tool. My side is a small TypeScript plugin that registers exactly three whitelisted, read-only tools, sift_check_claim, sift_extract_claims and sift_search_pubmed, each a thin proxy to the core. OpenClaw runs the plumbing and never touches the verdict; that still happens downstream, in the deterministic engine.

The report

What comes back when you hit check

This is the whole report for "creatine supplementation causes hair loss," exactly as it renders, top to bottom. The verdict leads, with confidence, evidence quality and directness as separate chips so none of them can quietly stand in for the others. Then the bottom line, the caveats, and the field I care about most (what would change this answer), which is stored in the report, not improvised, because a fact-checker that can’t tell you what would move it isn’t being honest. Below that sit the evidence pyramid and the direction chart, the split between mechanism and human outcome, the claim broken into checkable subclaims each with its own verdict, and finally the study cards, every one carrying its weight in the verdict and a live link out to PubMed.

localhost:3000/report · the running build
Sift evidence report for the claim that creatine supplementation causes hair loss: verdict Weak evidence, bottom line, what would change this answer, evidence pyramid, direction chart, mechanism versus outcome panels, subclaims, and seven study cards with PubMed links.
One page, built from the structured report object. The honest parts are the ones I’m proudest of: the retrieval noise is shown and labelled "not directly relevant" rather than swept away, and the verdict reads "Weak evidence," not "false."

I picked this claim on purpose, because it’s the perfect stress test. The creatine-and-hair-loss scare traces almost entirely back to one three-week study in twenty rugby players that measured a rise in a hormone marker, not actual hair loss, and was never replicated. The report says precisely that: mechanistic evidence weak, human-outcome evidence weak, a single small trial, and a "what would change this answer" list that asks for the study nobody has run yet. That shape, a real mechanism doing the work of a conclusion it hasn’t earned, is how most viral health claims are built, and I designed the report to make it visible.

The verdicts

Why I never let it say true or false

A binary verdict is the fastest way for a fact-checker to lose people on health topics, because the honest answer is almost always "there’s something real here, but the claim runs well past it." So I never let Sift say true or false. It chooses from eight cautious labels instead, and shows confidence and evidence quality right next to the label, so the uncertainty is part of the answer rather than something the tool quietly hides to sound sure.

VerdictWhat it means
SupportedGood evidence broadly supports the claim.
Likely supportedEvidence leans supportive but has real limitations.
Mixed evidenceFindings differ or depend heavily on context.
OverstatedSome truth exists, but the claim is exaggerated.
Weak evidenceThe evidence is limited or low quality.
UnsupportedEvidence does not support the claim.
ContradictedStronger evidence points against the claim.
Not enough evidenceToo little direct evidence to conclude.
Breaking a claim apart

A broad claim is really five questions in a coat

"Seed oils are bad for you" can’t be answered as written, because it’s hiding four or five different questions that have four or five different answers. So before Sift checks anything, the parser splits the claim into subclaims and checks each one on its own terms. It’s the same machinery that lets a single influencer sentence unpack into several linked reports instead of one mushy verdict.

flowchart TD
  A["Seed oils are bad for you"] --> B["Do seed oils cause cancer?<br/>Unsupported as a broad claim"]
  A --> C["Do seed oils cause inflammation?<br/>Mixed / context-dependent"]
  A --> D["Are repeatedly heated oils harmful?<br/>More plausible concern"]
  A --> E["Are seed oils worse than butter?<br/>Usually not supported"]
  A --> F["Are all seed oils the same?<br/>No, the category is too broad"]
One claim in, five checkable questions out, each carrying its own verdict.
The evidence store

Where the evidence lives, and how I rank it

A claim checker is only ever as honest as what it can pull up, so this is the part I spent the most time on. My bet was that the text search is the replaceable piece and the structure around each study is the durable one. So I hid the store behind one narrow interface (available, store_source, search → list[Source]) and nothing downstream knows or cares which backend answered. That’s what lets me swap the whole retrieval engine without touching the pipeline, the verdict, or the UI.

The unit I store isn’t a passage of text, it’s a typed Source: the title and citation, yes, but also the study type (systematic review, RCT, cohort, animal, in-vitro), the human relevance, the directness to the claim, the evidence direction (supports, mixed, contradicts, not relevant), the weight it earns in the verdict, the main finding and its limitations, each as its own field. The backend indexes the study’s text for search but also tucks away a lossless JSON copy of the object, so retrieval hands the pipeline fully-typed studies rather than paragraphs it has to re-read. That typing is the reason the report can say "mechanism only, low weight" and mean the exact same thing every single time.

Underneath that interface I run two backends. The default is SQLite FTS5, ranked by BM25 (no install, no keys, deterministic), which is the right call for a portfolio-stage build because anyone can clone the repo and it just works. The optional one, behind an EVIDENCE_BACKEND=gbrain flag, is a hybrid vector-and-keyword semantic store I ported from an earlier Dataclaw project, and it catches the case plain keyword search misses: "does linoleic acid raise cancer risk" and "are seed oils carcinogenic" are the same question wearing different words. Because both backends satisfy the identical interface, moving from keyword to semantic retrieval is a config change, not a rewrite.

Retrieval only finds candidates; deciding what each one is worth is a separate job, and that part is fully coded rather than aspirational. A claim taxonomy (causal-risk, benefit, superiority, mechanism, absolutist) sets what a good answer to that kind of claim should even look like. A study-design hierarchy (the evidence pyramid) assigns each design a quality value, from a meta-analysis at 1.00 and an RCT at 0.85 down through cohort at 0.60 to animal at 0.20 and in-vitro at 0.10. The ranker folds that together with human relevance, recency, sample size, directness and risk of bias into one transparent 0..1 score per study, and that score becomes the High, Medium or Low weight you see on each card. Every number in the report walks back to it.

I’ll be straight about what’s built and what isn’t. What runs today is a typed store with pluggable keyword or semantic retrieval and a ranking taxonomy that’s real and working. Where it’s headed is ontology-guided GraphRAG: storing each study as an exposure acting on an outcome, so retrieval can walk from a claim to evidence that never shared its vocabulary. I drew the interface the way I did specifically so that becomes the next backend, not the next rewrite.

LayerJobStatus
Typed storeEach study a structured Source: design, human relevance, directness, direction, weight, limitationsBuilt
RetrievalSQLite FTS5 / BM25 by default; optional gbrain hybrid vector + keyword, same interfaceBuilt, pluggable
Ranking taxonomyClaim types, evidence pyramid, transparent 0..1 per-study scoreBuilt
Ontology graphExposure–outcome graph traversal for vocabulary-independent retrievalNext backend
Under the hood

The pipeline, and what I don’t let the model do

A claim runs through a chain of deliberately narrow steps: parse it, decompose it, build the queries, retrieve from the store, classify each study, rank it, synthesize the direction, map that to a verdict, and run the safety layer. But the decision I care about most isn’t any one of those steps: it’s what I refuse to let the language model touch.

flowchart TD
  A["User claim<br/>web or Telegram"] --> B["Claim parser"]
  B --> C{"One claim<br/>or many?"}
  C -->|multi-claim post| D["Decomposer<br/>split into subclaims"]
  C -->|single| E["Normalize claim"]
  D --> E
  E --> F["Query builder<br/>generate PubMed queries"]
  F --> G["Evidence store<br/>seeded + live PubMed if thin"]
  G --> H["Study classifier<br/>type / human vs animal / directness"]
  H --> I["Evidence ranker<br/>weighted score per study"]
  I --> J["Evidence synthesizer<br/>direction + contradictions"]
  J --> K["Verdict engine<br/>DETERMINISTIC over structured evidence"]
  K --> L{"Safety gate<br/>high-risk topic?"}
  L -->|yes| M["Escalate caveats + clinician referral"]
  L -->|no| N["Standard disclaimer"]
  M --> O["Report generator"]
  N --> O
  O --> P["Web report"]
  O --> Q["Telegram short verdict + link"]
The verdict is computed, not narrated. The model only parses the messy input and writes the plain-English layer at the very end.

The model never grades evidence from memory, and I keep it boxed into three jobs, all of which I can audit. It parses messy input into testable claims. It labels each retrieved study’s stance (supports, contradicts, mixed, or not relevant), judging only from that study’s own finding text. And it writes the plain-English summary, but only after the verdict is already fixed. The verdict itself is a pure function over the direction tallies, the evidence quality, the directness and the claim type. Which means swapping the model can only ever move a verdict by relabelling an individual study, something you can check against the abstract yourself, and it can never cite a paper that wasn’t actually retrieved. That, more than anything else, is how I keep it from hallucinating a confident answer.

How a study earns its weight

Scoring factorWeight
Study design quality30%
Human relevance20%
Consistency across studies15%
Recency10%
Sample size / power10%
Directness to the claim10%
Risk of bias / conflicts5%

The verdict engine reads and writes one structured object per claim, and it’s worth looking at the last field. "What would change the answer" is stored right alongside the verdict, generated as part of the same pass rather than tacked on, so the report always carries its own falsification conditions.

claim_report.jsonJSON
{
  "primary_claim": "Vegetable oils cause cancer",
  "verdict": "unsupported_as_broad_claim",
  "confidence": "moderate",
  "evidence_quality": "moderate",
  "directness": "medium",
  "evidence_breakdown": { "systematic_reviews": 4, "rcts": 1,
                          "cohort": 12, "animal_or_mechanistic": 18 },
  "evidence_direction": { "supports": 2, "mixed": 7,
                          "contradicts": 10, "not_relevant": 4 },
  "caveats": ["Repeatedly reused frying oil is a separate question",
               "Oil types differ; the category is too broad"],
  "what_would_change_the_answer": [
    "Large prospective human studies showing dose-response harm",
    "RCTs with disease-relevant outcomes by oil type"
  ]
}

The stack, one job each

OpenClaw runtime

The local-first, messaging-first agent runtime. It runs the Telegram gateway, conversational state, and the agent loop, self-hosted on hardware I control.

Pluggable evidence store

SQLite FTS5 (BM25) by default, optional gbrain hybrid vector + keyword, behind one interface. Each study stored as a typed object, not a blob.

Python + FastAPI core

The evidence brain: parse, retrieve, classify, rank, verdict, safety, report. Every decision lives here; the surfaces are thin clients.

Next.js + TypeScript web

The home page, report, and library, calling the core over HTTP. The Telegram bridge is a separate TypeScript OpenClaw plugin.

Seeded + live PubMed

~25 common claims seeded via the NCBI E-utilities ingest CLI; novel claims with thin coverage retrieve live from PubMed at query time and cache back.

Provider-agnostic LLM layer

Parsing and summary sit behind one interface, so the model vendor can change without touching the pipeline.

Deterministic verdict engine

Plain scoring code over the structured fields. No model call, and no OpenClaw agent, decides a verdict, which is what keeps citations honest.

The hard parts

The bits that fought back

The first thing I got wrong was assuming quality and directness were the same axis. A rigorous meta-analysis on vitamin retention in pressure-cooked vegetables is genuinely high quality, and it still barely answers "is pressure cooking bad for you." Until I split those into two separate scores, both computed and both shown, the report kept quietly overweighting a precise answer to the wrong question.

Then there was mechanism inflation, which turned out to be the whole game. Almost every viral health claim points at a real mechanism (oxidation, a hormone marker, an inflammatory pathway), and the mechanism usually is real. The hard part was making "this mechanism exists in a dish" and "it matters at the amount you eat at dinner" land on different verdicts, so a rat study about overheated oil never gets to read as proof your cooking causes cancer.

Pulling claims apart was harder than it looked too. Influencer posts weld four claims into one breathless sentence, and the extractor had to separate them without inventing a claim that wasn’t there and without dropping one that was. That’s a much narrower target than "just split on the conjunctions," and it took real iteration to hit.

I was also surprised by how much work the tone took. Writing verdict language that informs without frightening ate as many revisions as the ranking did. "Overstated" has to read as a fair, slightly boring assessment, not a scold, because the moment it feels like a scold, people stop trusting the tool on exactly the topics that scared them enough to check.

And then the trade I made on the database itself. Seeding a snapshot kept the demo deterministic and honest, with no hallucinated citations, at the cost of breadth. I made that call on purpose for a portfolio-stage build, and I made it survivable: because the store sits behind one interface, the default SQLite/BM25 backend, the optional gbrain semantic one, and live PubMed at query time are all the same shape to the pipeline. That’s why live retrieval already backfills a novel claim when the local shelf is thin, without anything downstream needing to change.

Honest scope

What it can and cannot do

Can, today

  • Check nutrition, fitness, supplement, and sleep claims end to end.
  • Decompose a bundled multi-claim post into separate checkable subclaims.
  • Retrieve live from PubMed when a novel claim has thin local coverage, and cache it back.
  • Show its evidence, its direction, and its uncertainty, never a bare label.
  • Answer from Telegram and link to the full web report; save reports to revisit.
  • Run its own eval harness: verdict-in-band accuracy and high-risk recall against a gold set.

Cannot, on purpose or not yet

  • Cover PubMed broadly yet; the seeded store is ~25 common claims, with live retrieval only on the edges.
  • Diagnose, or advise on medication, pregnancy, pediatric, or eating-disorder topics.
  • Read screenshots, URLs, or video transcripts.
  • Personalize to your labs, history, or medical conditions.
  • Give a plain true or false; that is a design choice, not a gap.
The safety layer

Where I make it step back on purpose

A detector watches for the topics where a breezy evidence summary could actually do harm (pregnancy, children, cancer, kidney and liver disease, medication interactions, eating disorders, severe symptoms), and when it sees one it escalates the caveats or refuses to give treatment advice at all. Sift will never tell someone to stop a medication or hand out a personalized plan, because it has no tool that could. And every report carries the same disclaimer whether the claim was about frying oil or something far more serious, so the line never blurs.

This is an evidence summary, not medical advice. For medical conditions, medications, pregnancy, eating disorders, severe symptoms, or personalized treatment decisions, consult a qualified clinician.

I never wanted to tell anyone what to eat. I wanted to make "what does the evidence actually say" a one-minute question instead of an afternoon in PubMed.

python + fastapi core · openclaw telegram bridge · runs on my machine ← Back to the lab