Model, system, human: get the boundary right
Thirteen years of shipping AI, four arenas, the same discipline in each. Every card links to work that shows it.
Agentic AI platforms
Led product for enterprise agent platforms: self-service builders, agent libraries, RAG workflows, evaluation, governed deployment.
See it in practice →AI product operating models
Built the operating model that moved AI from custom delivery to repeatable product: discovery through release, metrics, feedback loops.
See it in practice →Personalization & decisioning
Led martech product lines end to end: recommendations, next-best actions, loyalty and lifecycle campaigns; AI content scaled to 8M+ personalized emails a week for QSR, airline, and fashion retail brands.
See it in practice →Responsible AI in production
Productized the controls for regulated environments: grounded outputs, human review, audit trails, PII controls, evaluation loops.
See it in practice →Thirteen years, one through-line
From data scientist to AI product leader: the analytics-to-product bridge the whole way. Products and public client proof are named where they are public; confidential client details stay private inside the case studies attached to each role.
Independent AI Product Advisor
- Advising early-stage teams across the AI product stack: algorithm and data modeling through product, UX, and agentic architecture.
- Hands-on as a builder: prototyping LLM and multi-agent workflows, validating model and eval choices, shaping 0→1 roadmaps.
Associate Principal & Head of Product Management
- Built out ZS's product management function, leveraging the firm's deep analytics and consulting muscle to productize solutions for the retail, F&B, and airlines sectors: product-market fit, a product shell to rally teams and customers around, and the discipline to make it real, then scaled it into a large cross-functional org across product, engineering, data science, design, and delivery.
- Stayed hands-on with the technical core: prototyped agentic workflows, defined how agents were evaluated before they reached production, and applied a data scientist's judgment on data, retrieval, and failure modes to product decisions.
- Drove product strategy, roadmap, GTM alignment, and cross-functional execution across three commercial product lines: a patented martech personalization engine for marketing and loyalty teams (Personalize.AI), an agentic AI platform (Max.AI), and a regulated-content supply chain (ZAIDYN Content); the personalization work contributed to ZS being named a Forrester Leader in Customer Analytics Services (2025).
- Directed the shift from fragmented custom ML to a reusable agentic + GenAI platform architecture, reducing duplicated delivery effort and enabling repeatable deployment of specialized agents across regulated and commercial use cases.
- As part of the leadership team, helped introduce the forward-deployed engineering model and shape the portfolio investment case, and partnered closely with customer success, FDEs, and clients through pilot deployments to make sure solutions succeeded in production.
Products I led
Selected public proof
AWS Events
AWS re:Invent 2025
AWS Events
AWS re:Invent 2024
AWS Events
AWS re:Invent 2023
ZS & Cerebras Partner on Agentic AI
ZS Named a Leader: Forrester Wave
Personalize.AI Earns U.S. Patent
ZS Launches ZAIDYN, MaxAI Self-Serve AI
Triumph Group Partners with ZS
LinkedIn
Panera Bread Runs on Personalize.AI
LinkedIn
AI & Agentic Summit: Singapore & Japan
IMF
Myelo · IMF
Director, AI/ML Product Studio
- Founding-team role: product lead, commercial director, and solution architect at once, building an AI personalization studio from concept to production.
- Shipped Alice (an enterprise virtual concierge for hospitality) and a retail personalization product distributed via Shopify and Magento.
- The product, IP, and core team carried directly into ZS in 2021 to become the foundation of Personalize.AI.
Selected work
Data Scientist & Engagement Lead
- First-50 employee at an early-stage analytics firm, hands-on IC and internal builder simultaneously.
- Built end-to-end as sole data scientist: designed, coded, and deployed the demand-forecasting model and pipeline for a global FMCG client: feature engineering, time-series regression, and client-facing reporting layer, no handoffs.
Selected work
Data Scientist
- Where it started: individual contributor practitioner, writing models and shipping them to production across concurrent client engagements.
- Long-term Germany onsite with a global sportswear brand: personally designed and deployed category-affinity, purchase-propensity, and product-affinity models powering on-site recommendations, each validated in-market via A/B tests I designed and ran.
- Sole data scientist on a dynamic pricing engine for one of the largest US secondary ticketing marketplaces: demand modeling, price sensitivity, and real-time recommendation output built and shipped end-to-end.
Selected work
Technical fluency
Deep product fluency
Hands-on / prototyping

A product leader who started as a data scientist.
I'm Nandini, originally from Ooty, based in Bengaluru. I started as a data scientist, and it still shapes how I lead: I'm comfortable with the mechanics underneath AI products (data, retrieval, evaluation, experimentation, failure modes), but my work today is product leadership. Deciding what to build, how it should work, how it's evaluated, how teams ship it, and how customers adopt it.
I'm an ambivert: happiest deep in a problem with a small group of sharp people, and just as energized kicking off a big room at a conference. Off the clock, I trade dashboards for dirt and paws: a balcony that's slowly become a small jungle, and a cat with strong opinions who's the closest thing I have to a stakeholder review.
Three rules the builds keep proving
Decide what the model may decide
I split every workflow into what the model generates, what the system computes, and what a human inspects. Most AI product failures are boundary failures, not model failures.
No receipt, no claim
For high-stakes or subjective output I design for traceability: citations, pins, scores, review states, visible failure modes. The verdict is computed, not narrated.
Enter the workflow, don't replace it
The best AI feature joins a workflow people already run (a chat thread, a template, a review step) instead of asking them to learn a new system.
Small builds, real product questions
I use small builds to test product questions I can't answer from demos alone. Some became working tools, some are deliberately odd. The point is the same every time: build the workflow end to end, find where the model breaks, decide what should be deterministic, and turn the lesson into product judgment. Right now that's three threads: evidence and grounding (Sift, Brand Mood Ring), automation inside constraints (the PowerPoint generator), and generative media (Blankie, Luna Whisk).
Sift: a fact-checker for the health claims you hear
Can checking a health claim feel as easy as forwarding a message? Paste "seed oils cause cancer" and Sift decomposes the claim, pulls the PubMed evidence, scores study quality and directness separately, and computes a cautious verdict, on the web or straight from Telegram. The model parses and explains; it never gets to be the judge.
Read the build →Brand Mood Ring
Can AI critique brand feel without going vague, mean, or ungrounded? A playful multimodal scanner that pins every observation to a quoted line, a visual region, or a palette value.
Read the build →Oh, I Can't Find My Blankie!
Could early image models hold one character steady across a whole story? A self-published picture book made right as DALL-E made that almost possible: story, art, and a lot of redos.
Read the story →Inter-Dimensional Adventures of Luna Whisk
How much continuity can a film keep when every shot is generated? A 48-hour AI filmmaking sprint for RunwayML's first GEN:48 challenge: one brief, one deadline, a finalist film.
Watch the film → Constrained automation · 2025A PowerPoint generator, before the models shipped one
Can agents generate useful slides without breaking the template? Extract the deck into a layout contract, hand the model slots and character budgets, reject overflow.
See how it worked →Essays on practical AI
Long-form pieces on how AI products work in production, published on LinkedIn and Medium.
Short posts
Shorter takes worth sharing.

Oh, I Can't Find My Blankie!
AI images went from lucky screenshots to finished work
In 2022, image models were new enough that the honest question wasn't "is this pretty" but "can it carry something finished": a dozen pages that hang together, not one lucky screenshot. Two years later that question sounds quaint: the tools reset every few months and quietly swallow whole workflows. Blankie is a snapshot from inside that shift, and the clearest way I've found to show how far it's moved.
A picture book is a brutal little benchmark
A children's book is a deceptively hard thing to point an image model at: a dozen-plus spreads, and the same teddy bear has to read as the same bear on every one, from every angle, in the same hand. Back then that one requirement (character consistency) was precisely what the technology was worst at, which is exactly why it was worth trying. The writing was the gentle part: plain couplets a sleepy kid could follow, like "But oh no! Where's his blankie red?", drafted with ChatGPT as a sparring partner for rhythm. Everything hard was in the pictures.
These were pure text-to-image systems with no memory between generations: each image an independent roll of the dice, with no notion of this character that carried from one prompt to the next. So the whole aesthetic had to be re-specified, in full, every single time, which made the prompt the only real lever. Ours lived as one long, exact paragraph of illustration rules (ink outline, no shadows, no hatching, a light watercolour palette, a bluish wash) where word order mattered. You generated in batches, cherry-picked the survivors, and leaned on the workarounds of the day: fixed seeds, image-to-image, inpainting, pose conditioning.
A whole casting call
None of the crew, bear, bunny, owl, cat, dog, started as a fixed design. This is a slice of the rounds across both DALL·E and Midjourney it took just to find a look worth committing to, well before keeping it consistent was even the problem.
The turning point was Midjourney's character-reference feature, --cref. For the first time you could hand the model an approved portrait and say, in effect, this one, tuning how hard it listened with --cw. It was the single most useful thing in the whole project. We settled on a Teddy everyone agreed on, then pointed every subsequent scene back at that reference.
Teddy, locked
This portrait became the --cref for the rest of the book. The jacket, the proportions, the expression, all of it had to survive into every other page, at every other angle, for months.
The prompt that made him "Teddy"
A highly simplified, hand-drawn ink illustration of a small brown teddy bear named Teddy, wearing an orange zip-up jacket and denim shorts. The style is minimalistic, with minimal details and no shadows, no hatching, outlined in black ink for a crisp, clean appearance. The colors used are very light, watercolour, and pastel, emphasizing a gentle aesthetic. Forced perspective, Teddy mid-jump, arms out, looking straight at camera, a few motion marks trailing his feet. --v 6.0
From there, every scene reused that same style paragraph with --cref aimed at the portrait. It helped. And it still wasn't enough.
Still four different bears
This is the honest limit of that era, in one batch: identical prompt, identical reference, identical weight. One reads almost photoreal, one is a loose painterly sketch, two are close but not twins. The reference narrowed the range; it never closed it. Every batch still ended with a human picking the one that belonged in the book.
Even the keepers weren't done. Each needed real Photoshop work: a stray extra finger, lighting matched across spreads, text composited into the art, print files prepped. The "AI-generated book" was, in practice, a lot of manual craft stitched around the generations. And the sharpest editor in the whole process wasn't a model at all: Priya went through the draft page by page, flagged that naming every character right before bed was one too many for a drowsy toddler so that page got cut, then read the whole thing aloud to her own kid and reported back on what landed.
The finished pages
All that casting and re-generating, in service of three quiet spreads: Teddy ready for bed, Teddy realizing blankie's gone, Teddy looking here, then there. It went out as a real, self-published, if-you-lose-it-you-owe-me-a-copy book, with a calming narrated version following on YouTube much later.
The gap that closed in three years
Almost every hard-won trick in the book has since been absorbed into the models themselves. The re-typed style paragraph, the --cref reference, the seed-locking, the batch-and-cherry-pick, the inpainting, the Photoshop salvage. That entire apparatus existed to paper over one gap: the model couldn't hold a character in its head.
Now it largely can. Multimodal, in-context generation lets you hand a model your character and your instructions in the same breath, keep it consistent across a whole set without a special flag, and, most tellingly, edit in place by conversation: "same bear, now on a swing, keep the jacket," instead of rolling the dice again. That moves where the skill lives: when consistency is nearly free, taste and art direction matter more than prompt-craft. Re-made today, Blankie would be closer to an afternoon than the months it took.
Made for bedtimes, not a portfolio
This was never a portfolio exercise. It got made to be read to actual kids at actual bedtimes, argued over on WhatsApp threads, and revised because a real toddler lost interest on the wrong page. It's easy to track this shift from release notes and demos; Blankie is my argument for the opposite: you only really feel how far the tools have moved by trying to make something finished with them, then trying again a year later.
The book is finished. The field it came out of isn't, and that's the part I can't stop watching.
Inter-Dimensional Adventures of Luna Whisk
A short film made with a few friends in 48 hours flat, for the first-ever edition of RunwayML's GEN:48 challenge: one location, one character, and one event handed to every team at the same moment, with a hard deadline and no extensions. Ours made the finalist list.
Watch it on RunwayMLAI video moves faster than almost anything I've worked on
Text-to-video has gone from a party trick to something you can almost put in front of an audience, in about two years. Whole categories of work (shorts, ads, explainers, music videos) now get roughed out by people who'd never go near a render farm, and the state of the art turns over on a scale of months. I've been making things inside that churn since close to the beginning; the first one that felt real was a 48-hour film called Luna Whisk.
Luna Whisk, start to finish in a weekend
GEN:48 was RunwayML's first 48-hour film challenge. Every team gets the same three ingredients at the same moment (a location, a character, an event), then races a hard deadline with no extensions. Ours became an interdimensional cat named Luna Whisk, made with a few friends over one weekend.
The pipeline was whatever we could stitch together. DALL·E generated the interdimensional worlds. Runway's image generator handled character shots, fed reference images to keep Luna on-model. Its Gen-Turbo model turned those stills into motion. An original score came from Suno, sound effects from the Epidemic Sound library the contest handed participants, and narration from an AI-cloned voice in ElevenLabs. A friend who does this for a living wired the tools together, knew which warped takes to throw out, and did the final color and edit, exporting minutes before the cutoff.
This edition predated most of today's tooling, so those five tools were, between them, close to the entire stack. And the hard part was never the look. It was continuity. Nothing enforced that the cat in one shot was the cat in the next, so "the same cat" meant close enough, picked by eye.
Six tools in, one film out
Every box below was a separate product, and every arrow was a hand-off we made by hand. That topology is why continuity was the hard part: nothing downstream knew what the step before it had produced.
What "a cat character" got back
A few of the early passes: a shutterstock-style mascot in a hooded robe, a photoreal ginger stray on a lily pad, a painterly kitten curled in a teacup. None of them were the white cat with amber eyes and a glowing third-eye gem that eventually shipped. Landing on a reference worth locking in took more rounds than getting a good scene once we had one.
Luna Whisk, mid-scene
One of the reference-anchored character shots from Runway: the same white cat with amber eyes and the glowing third-eye gem, dropped into a new backdrop and still recognizably herself.
Same character, different pass
This one came out of a different generation pass, and it shows: a striped tail and softer markings, next to the clean white-and-orange coat in the earlier stills. Nothing in the pipeline enforced one source-of-truth reference across every scene, so "the same cat" meant close enough, picked by eye, more often than it meant identical.
Then, the finalist email
From stitched clips to something closer to a world model
The reason all that stitching was necessary is worth naming. Runway in 2023 was really an image model taught to move: each shot was denoised on its own, with no memory of the one before it, so it had no concept of the same cat, only a plausible next frame. Almost everything interesting since has come from dropping that one assumption.
The current wave of models treats a clip as a single object in space and time instead of guessing frame by frame, which is most of the gap between the old "AI melt" and a shot that holds a coherent character and light source across ten seconds. References are a first-class input now rather than a folder of screenshots; models that score their own audio fold the Suno-and-ElevenLabs step back into the render; the five-tool pipeline is collapsing toward one. It's nowhere near solved (long-range consistency and directing a shot are still open), but "keep this one cat consistent" went from the whole challenge to almost a checkbox in two years.
What moved, dimension by dimension
| Dimension | 2023 · what I worked with | The shift since |
|---|---|---|
| Architecture | Image diffusion, one clip at a time | A spacetime diffusion transformer: the clip modelled as one object, not guessed frame by frame |
| Temporal unit | 2–4s clips, each generated alone | 10s+ coherent shots in a single pass |
| Character | Reference images + pick-by-eye | References as a first-class, built-in input |
| Cross-shot continuity | None enforced; held by a human | Emerging shot & scene memory |
| Audio | Suno + ElevenLabs, stitched in after | Native audio & dialogue in the same render |
| Complex motion | Warps and morphs, "AI melt" | Rough physics & object permanence |
| Pipeline | Five tools + manual glue | Collapsing toward one or two |
Reference points for the right column: Sora (spacetime patches, 2024), Runway Gen-3/Gen-4, Google Veo, Kling, Luma Dream Machine.
Where a capability actually is, versus the demo
Here's the professional reason I build things end to end like this: it's how I learn where a capability is: not where the demo puts it, but where it breaks, what's still held together by people, and which of those seams the next model quietly absorbs. Reading the real frontier against the marketed one is most of what I bring to AI product work, and it doesn't come from a changelog.
The un-strategic reason is simpler: I find this genuinely amazing, and I'd rather stay close to the thing than only read about it. Luna Whisk is what came out of following that: a weekend with friends, made to see what would happen.
Luna promises endless interdimensional cuddles in return. 🐱
A PowerPoint generator, built the year before the models could
In early 2025, "make me a deck about X" was still an open gap. You could get an LLM to write slide copy, but what came back looked like AI slop. Centered text on a white background, none of the craft of a deck someone designed. So over a couple of weekends I built the opposite: not a model that designs slides, but a small crew of agents that fill ones a designer already got right. It worked end to end, rendering real decks with real content and visuals.
The bets, up front
| Product question | Can agents generate useful decks inside an existing design system? |
| User | A team that needs a decent first draft without wrecking the template. |
| Design bet | Don't ask the model to design the slide. Give it slots, budgets, and constraints. |
| System bet | Extract the .pptx into a layout contract, generate into that contract, reject anything that overflows. |
| What broke | The model treated character limits as vibes until the fit check made them measurable. |
| What I'd change | Keep the contract, swap the writer. The loop, not the prompt, is the part that aged well. |
A good slide is two things wearing one coat: a design (where the title sits, what the palette is, how many characters a box holds before it looks cramped) and the content that lands inside it. Models in early 2025 were strong at the second and hopeless at the first.
So I stopped asking them to do the first at all. I bought a professionally-designed template (a light modernist pitch deck, thirteen slides) and treated the design as fixed. The whole system had one job: read that deck into a spec precise enough that agents could pour new content into it without touching a single design decision. The template is the contract; the agents just honor it.
#3AEFCC, Arial Black titles). The dashed boxes are the two text slots the extractor found, each tagged with the exact character budget, pulled from the deck, that the agents have to write within.First I asked the model to read the deck. Then I stopped guessing.
The obvious first build was to dump every slide's shapes into GPT-4o and ask it to describe the template as JSON. It demoed fine and failed in practice: the model hallucinated placeholders that weren't in the deck and dropped ones that were. It was guessing at a file it couldn't see.
# v1 (rejected) — ask the model to infer the template from a shape dump
resp = openai.chat.completions.create(
model="gpt-4o",
response_format={"type": "json_object"},
messages=[
{"role": "system", "content": "You are an expert in slide design."},
{"role": "user", "content": prompt}, # raw shapes + instructions
],
temperature=0.7,
)
# Failure: invents placeholders not in the deck, ignores ones that are.
So I threw it out. A .pptx is just a zip of XML, so v2 reads the actual document instead of describing it: every text frame's position, font, size, weight, color, alignment, and bullet levels, pulled from the drawing-markup, with python-pptx as a fallback when the XML stays silent. No guessing; the spec now describes the deck that exists.
# v2 — read the drawing-markup directly (source of truth, not a guess)
NS = {"a": "http://schemas.openxmlformats.org/drawingml/2006/main"}
for p in txBody.findall(".//a:p", NS):
defRPr = p.find(".//a:pPr/a:defRPr", NS)
if defRPr is None: continue
if "sz" in defRPr.attrib: # size stored x100
font_size = float(defRPr.attrib["sz"]) / 100.0
latin = defRPr.find(".//a:latin", NS)
if latin is not None:
font_name = latin.attrib.get("typeface")
srgb = defRPr.find(".//a:solidFill/a:srgbClr", NS)
if srgb is not None:
text_color = "#" + srgb.attrib["val"]
Everything the extractor produces lands in one typed record per text frame. Get the schema right and the reading half and the writing half can be built and swapped independently. They only ever agree on this shape, and max_chars is the line that makes the contract enforceable:
@dataclass
class TextFrame:
id: str # unique per frame (uuid4[:8])
placeholder_type: str = None # CENTER_TITLE, SUBTITLE, BODY, ...
text: str = ""
left: float = 0.0; top: float = 0.0; width: float = 0.0; height: float = 0.0
max_chars: int = 0 # the constraint the agents must respect
font_name: str = None; font_size: float = 0.0; font_bold: bool = False
text_color: str = None; horizontal_alignment: str = None
bullet_points: list = field(default_factory=list)
level: int = 0 # 0 title, 1 subtitle, 2 body
Built from six pieces, each doing one job
python-pptx
The friendly path into OOXML: iterate slides, shapes, placeholders; read geometry in EMU, convert to inches, and at the end write filled copy back in place.
xml.etree.ElementTree
The escape hatch. When the high-level API returns None for inherited styling, I drop to the raw a:defRPr markup and read it myself.
dataclasses
Typed TextFrame / SlideTemplate as the spec schema; asdict() serializes straight to JSON. The schema is the interface between extractor and agents.
collections.Counter
Real decks are messy: runs disagree on font. Counter resolves the dominant font, size, and color per box.
openai · gpt-4o
The engine behind the planner and content agents, in json_object mode, alongside the rejected v1 kept in the repo as a documented dead-end.
uuid
Short ids per frame so an agent can address one specific slot (fill frame b5ee39cd) without positional ambiguity.
flowchart LR
A["Designed .pptx<br/>(bought template)"] --> B["PPTTemplateExtractor<br/>python-pptx + raw OOXML"]
A --> C["ThemeExtractor<br/>colors / fonts / master styles"]
B --> D["template_data.json<br/>per-slide spec + char budgets"]
C --> E["theme_data.json"]
D --> F{{"The contract"}}
E --> F
F --> G["Agent loop<br/>plan / write / fit-check"]
G --> H["Rendered .pptx<br/>design untouched"]
Inheritance, and giving every box a budget
Most slides don't state their own formatting. A title box inherits its font from the layout, which inherits from the master, which falls back to the theme. Read the shape alone and you get None for nearly everything, so the extractor walks the same cascade PowerPoint resolves at render time, stopping at the first level that answers.
flowchart TD
S["Shape run properties"] -->|null?| L["Layout placeholder"]
L -->|null?| LD["Layout default font"]
LD -->|null?| M["Master default font"]
M -->|null?| T["Theme font scheme"]
T -->|null?| F["Hardcoded fallback<br/>Arial 12pt, #000"]
S -.first hit.-> R{{"Effective formatting"}}
L -.-> R
LD -.-> R
M -.-> R
T -.-> R
F -.-> R
The second part is the one that makes the whole thing enforceable. For each frame the tool estimates how much copy fits from the box's real size and font, and writes that number into the spec as a hard limit. Downstream, the content agent is told "this subtitle gets ~324 characters," not "write a subtitle." That budget is what later lets a fit-check reject a slide that overflows.
def _estimate_max_chars(self, width, height, font_size):
# avg glyph advance ~= 60% of point size (heuristic, not true metrics)
chars_per_inch = 72 / (font_size * 0.6)
chars_per_line = max(int(width * chars_per_inch), 10)
line_height = font_size * 1.2 / 72 # points to inches
max_lines = max(int(height / line_height), 1)
return chars_per_line * max_lines # the box's copy budget
A small crew of agents, one feedback loop
With the contract in hand, generation stopped being a single prompt and became a loop of narrow agents, each with one responsibility. Give it a topic and the spec, and:
A planner chooses which template slides to use and orders them into a story. A content agent writes copy into each slot, one frame at a time, working against that frame's max_chars. A visual agent generates the charts and imagery, reading from the same theme palette the extractor pulled so nothing clashes. Then a fit-check measures every box against its budget, and anything over goes straight back to the content agent for a tighter rewrite. Only once every slide passes does the render agent write the filled spec back into a real .pptx through python-pptx, editing runs in place so the design survives untouched.
flowchart TD
U["Topic + template spec"] --> P["Planner agent<br/>choose slides, order the story"]
P --> C["Content agent<br/>write copy into each slot"]
C --> V["Visual agent<br/>chart / image in theme palette"]
V --> K{"Fit check<br/>every box within max_chars?"}
K -->|over budget| C
K -->|passes| Rn["Render agent<br/>python-pptx writes .pptx"]
Rn --> D["Finished deck<br/>design untouched"]
What fought back
Getting a model to respect a hard limit. An LLM treats "≤324 characters" as a polite suggestion. No amount of prompt-tuning made it reliable. The fix was structural, not verbal: generate, measure, and bounce anything over budget back for a rewrite. The fit-check loop is what turned a suggestion into a constraint, and it's the piece that made everything else trustworthy.
Rendering back without wrecking the design. Writing generated copy into a slide is deceptively easy to get wrong: replace a text frame wholesale and you lose the exact fonts, colors, and spacing the template existed to provide. The render agent had to edit runs in place, changing the words while leaving every formatting attribute the extractor had so carefully resolved.
Visuals that looked like they belonged. A perfectly-written slide still looks stitched together if the chart shows up in stock blue. Feeding the extracted accent palette into the visual agent kept generated graphics on-theme: small thing, big difference in whether a deck reads as designed or assembled.
The budget is still a heuristic. The character estimate approximates glyph width; it doesn't know an i from a W, so it's occasionally off. I shipped it anyway, because the fit-check downstream caught the misses, an imperfect budget was good enough to build a reliable loop on. Real per-typeface metrics would tighten it, and that's the first thing I'd redo.
The models shipped it, and got the framing right
I got it working end to end: a topic in, and the agents planned the deck, wrote to every box's budget, generated visuals in the theme palette, and rendered a real .pptx with the design intact. A few months after I had it producing decks, Claude and Gemini shipped native deck and document generation, and any reason to keep building a bespoke pipeline evaporated. That's the right outcome: if a frontier model does the whole job well, a weekend project shouldn't compete.
But it was satisfying to watch, because the thing they got right is the thing this was built around: quality comes from working within a strong design system, not from asking a model to conjure one each time. What stays with me isn't the code, it's the framing: find the taste-dependent half of a problem, make it a fixed input, and let agents do only the part they're reliable at, inside a loop that checks their work. That instinct outlives the prototype, which is the whole point of building small things at the edge of what's possible. You learn where the edge is with your hands on the file, not from the release notes.
Don't ask a model to design the slide. Ask a crew of agents to fill the one a designer already got right, and check their work.
Brand Mood Ring
Upload a creative. Find out what it really says.
Brand Mood Ring is a playful interface around a serious product question: can a multimodal system critique a marketing asset in a way that is grounded, useful, and emotionally legible? Drop in a poster, ad, landing-page screenshot, or a raw HTML email, and it returns a creative diagnosis: what it feels like, who it’s really for, and what it’s accidentally saying. The fun is in the surface. The discipline is underneath: every read has to point back to a quoted line, a visual region, or a palette value.
The bets, up front
| Product question | Can a multimodal system critique a marketing asset in a way that is grounded, useful, and emotionally legible? |
| User | Someone about to ship a campaign who half-suspects the asset is off but can't name why. |
| Design bet | A playful surface (creatures, roasts, a mood ring) so the critique lands as a read, not a report card. |
| System bet | A Python scout computes the receipts (OCR regions, palette values) before any model speaks; a referee bounces claims without evidence. |
| What broke | The critic sounded equally confident whether or not it was grounded. The referee pass exists because of that. |
| What I'd change | Add first-glance saliency signals so the 3-second read is measured, not asserted. |
I started with a dumb question
I wanted to know if a vision model could read vibe: not “this is an email with a button,” but “this is trying to be your funny friend, and mostly pulling it off.” One weekend turned into several, because the diagnoses kept coming back sharper (and funnier) than I expected.
What kept me building: the gap read turned out useful. Most creative isn’t bad, it’s mismatched. It intends one thing and signals another, and naming that gap in plain, slightly rude language turns out to be something a multimodal model is weirdly good at. Somewhere along the way the prompt chain grew a scout, a referee, and opinions about evidence, and became an actual app.
So I built a scanner with four reads
I fed it a Zomato email, Goblin Mode on
This is the app mid-diagnosis: Streamlit on top, a crew of Claude agents underneath. Left pane: the uploaded email with numbered evidence pins, drawn from OCR boxes a Python scout measured before any model said a word. Top right: the crew tracker, tagger to critic to referee to goblin, ticking green as each agent lands. The verdict: the Contraband Cookie Bandit.
That one scan fills six tabs. Each is a different read on the same pixels, and every screenshot below is real output, from this Zomato scan or the Olaplex one further down:
Then I pointed it at quiet luxury
To prove the ring reads vibe and not a checklist, I picked the Bandit’s exact opposite: an Olaplex nurture email. Premium, clinical, five thousand pixels of bond science. The ring named it the White-Coat Rapunzel, and the flaw it found is structural, not tonal: one email wired for two customer journeys, a welcome gift up top and a cart-recovery button right under the science. This is also the scan where the referee earned its keep, bouncing the critic’s first draft for ungrounded claims and forcing a redo.
How I built it: receipts first, opinions second
Ingest normalizes the input: images pass straight through; raw HTML gets rendered to pixels by headless Chrome and parsed for its DOM facts, so the model reads what the email says and what it looks like. Then a deterministic Python scout runs OCR and k-means clustering to produce the receipts: text blocks with ids and pixel boxes, a palette with exact percentages. Only after that does the crew speak: a tagger sorts, a critic interprets, a referee bounces any claim the scout can’t back up, and the goblin and remixer go last.
flowchart LR
U["Upload<br/>image or HTML"] --> IN["Ingest<br/>HTML rendered + DOM parsed"]
IN --> SC["Scout, plain Python<br/>OCR receipts + palette"]
SC --> TG["Tagger agent<br/>four-lens tags"]
TG --> CR["Critic agent<br/>vibe, gap, cliches"]
CR --> RF{"Referee<br/>every claim has a receipt?"}
RF -->|grounded| OUT["Structured JSON<br/>renders the six tabs"]
RF -->|bounced, redo once| CR
CR -.-> GB["Goblin agent<br/>roast + translation"]
CR -.-> RX["Remix agent<br/>dial-driven rewrites"]
GB -.-> OUT
RX -.-> OUT
flowchart LR D["Dial picked<br/>e.g. more wholesome"] --> RX["Remix agent<br/>drafts the rewrite"] RX --> CR["Critic agent<br/>re-scores the draft"] CR -->|voice drifted| RX CR -->|same DNA, new angle| SHIP["Rewrite ships"]
{
"creature": { "name": "the Contraband Cookie Bandit",
"chips": ["conspiratorial", "meme-fluent"] },
"scores": { "clarity": "High", "trust": "Medium", "distinctiveness": "High" },
"palette": [{ "color": "#1e1e21", "percentage": 36.5 }, "..."],
"evidence_pins": [{ "pin": 1, "receipt_id": "R1" }, "..."],
"referee": { "grounded": true }
}
How I test something with no answer key
Taste has no ground truth, so I test around it. The bar is three properties I can measure:
What I’m building next
It doesn’t ask whether your creative is good. It asks what your creative is making people believe.
Sift
Fact-check the health claims you hear.
A friend forwards a reel ("seed oils cause cancer") and you want to know if it’s true. The honest answer is buried in a research database built for scientists, not for someone standing in a kitchen. Sift is the layer I built in between. You paste the claim you heard, and it breaks the claim apart, pulls the evidence, weighs how good that evidence actually is, and hands back a calm, source-backed verdict, on the web or straight from Telegram. It runs end to end, and every screenshot on this page is the real thing.
The bets, up front
| Product question | Can checking a health claim feel as easy as forwarding a message? |
| User | Someone who just got sent a scary nutrition reel and wants one calm answer before acting on it. |
| Design bet | Start from the claim, not the paper. Keep the front door to one input. |
| System bet | The model parses and explains; the verdict is computed from structured, scored evidence. |
| What broke | Study quality and claim directness looked like one axis until the scoring forced them apart. |
| What I'd change | Grow the evidence graph past the seeded claims. The ontology, not the chat, is where this compounds. |
It started with a claim I couldn’t check fast enough
Someone sent me "creatine causes hair loss" as a fifteen-second clip. Checking it properly means knowing which of forty thousand PubMed hits actually speak to the question, which ran on people instead of cells, and which are strong enough to change what you’d do about it. That’s a researcher’s afternoon, and nobody watching the clip has one to spare.
Consensus and tools like it already proved people will use an evidence search engine as long as it talks plainly. I wanted something narrower and more opinionated: not a search box over papers, but a claim checker. You hand it the exact sentence you heard, and it does the translation, because people think in claims and PubMed thinks in papers. Everything Sift does lives in the gap between those two, and it always starts from the claim, never from the paper.
One box, and whatever you just heard
I kept the home page deliberately plain: one input, a few example claims so you’re not staring at an empty box, and a single button. Underneath it I put a wall of trending checks with their verdict chips already attached, so the vocabulary teaches itself before you type a word. You see "Unsupported" and "Likely supported" and "Overstated" sitting next to real claims and get the idea. There are no accounts and no dashboard. All the complexity I could push backward, I did; the page you land on doesn’t show any of it.
Catching the claim where you actually heard it
Health claims almost never arrive while you’re sitting at a browser. They show up mid-conversation, in the chat you already have open, so I put Sift there too. Send it a single claim and it replies with the verdict, the confidence, a couple of plain sentences, and a link to the full report if you want to go deeper. Paste a whole influencer rant and it does something I like more: it reads the blur, pulls out the separate claims hiding inside it, and asks which one you want checked. Both of these screenshots are the actual bot, on my phone.
The bot isn’t a one-off integration I hand-wired. It runs on OpenClaw, a local-first, messaging-first agent runtime, which means the Telegram gateway, the conversational state that remembers which claim you picked, and the agent loop all come for free and all stay on hardware I control, which matters more than usual for a health tool. My side is a small TypeScript plugin that registers exactly three whitelisted, read-only tools, sift_check_claim, sift_extract_claims and sift_search_pubmed, each a thin proxy to the core. OpenClaw runs the plumbing and never touches the verdict; that still happens downstream, in the deterministic engine.
What comes back when you hit check
This is the whole report for "creatine supplementation causes hair loss," exactly as it renders, top to bottom. The verdict leads, with confidence, evidence quality and directness as separate chips so none of them can quietly stand in for the others. Then the bottom line, the caveats, and the field I care about most (what would change this answer), which is stored in the report, not improvised, because a fact-checker that can’t tell you what would move it isn’t being honest. Below that sit the evidence pyramid and the direction chart, the split between mechanism and human outcome, the claim broken into checkable subclaims each with its own verdict, and finally the study cards, every one carrying its weight in the verdict and a live link out to PubMed.
I picked this claim on purpose, because it’s the perfect stress test. The creatine-and-hair-loss scare traces almost entirely back to one three-week study in twenty rugby players that measured a rise in a hormone marker, not actual hair loss, and was never replicated. The report says precisely that: mechanistic evidence weak, human-outcome evidence weak, a single small trial, and a "what would change this answer" list that asks for the study nobody has run yet. That shape, a real mechanism doing the work of a conclusion it hasn’t earned, is how most viral health claims are built, and I designed the report to make it visible.
Why I never let it say true or false
A binary verdict is the fastest way for a fact-checker to lose people on health topics, because the honest answer is almost always "there’s something real here, but the claim runs well past it." So I never let Sift say true or false. It chooses from eight cautious labels instead, and shows confidence and evidence quality right next to the label, so the uncertainty is part of the answer rather than something the tool quietly hides to sound sure.
| Verdict | What it means |
|---|---|
| Supported | Good evidence broadly supports the claim. |
| Likely supported | Evidence leans supportive but has real limitations. |
| Mixed evidence | Findings differ or depend heavily on context. |
| Overstated | Some truth exists, but the claim is exaggerated. |
| Weak evidence | The evidence is limited or low quality. |
| Unsupported | Evidence does not support the claim. |
| Contradicted | Stronger evidence points against the claim. |
| Not enough evidence | Too little direct evidence to conclude. |
A broad claim is really five questions in a coat
"Seed oils are bad for you" can’t be answered as written, because it’s hiding four or five different questions that have four or five different answers. So before Sift checks anything, the parser splits the claim into subclaims and checks each one on its own terms. It’s the same machinery that lets a single influencer sentence unpack into several linked reports instead of one mushy verdict.
flowchart TD A["Seed oils are bad for you"] --> B["Do seed oils cause cancer?<br/>Unsupported as a broad claim"] A --> C["Do seed oils cause inflammation?<br/>Mixed / context-dependent"] A --> D["Are repeatedly heated oils harmful?<br/>More plausible concern"] A --> E["Are seed oils worse than butter?<br/>Usually not supported"] A --> F["Are all seed oils the same?<br/>No, the category is too broad"]
Where the evidence lives, and how I rank it
A claim checker is only ever as honest as what it can pull up, so this is the part I spent the most time on. My bet was that the text search is the replaceable piece and the structure around each study is the durable one. So I hid the store behind one narrow interface (available, store_source, search → list[Source]) and nothing downstream knows or cares which backend answered. That’s what lets me swap the whole retrieval engine without touching the pipeline, the verdict, or the UI.
The unit I store isn’t a passage of text, it’s a typed Source: the title and citation, yes, but also the study type (systematic review, RCT, cohort, animal, in-vitro), the human relevance, the directness to the claim, the evidence direction (supports, mixed, contradicts, not relevant), the weight it earns in the verdict, the main finding and its limitations, each as its own field. The backend indexes the study’s text for search but also tucks away a lossless JSON copy of the object, so retrieval hands the pipeline fully-typed studies rather than paragraphs it has to re-read. That typing is the reason the report can say "mechanism only, low weight" and mean the exact same thing every single time.
Underneath that interface I run two backends. The default is SQLite FTS5, ranked by BM25 (no install, no keys, deterministic), which is the right call for a portfolio-stage build because anyone can clone the repo and it just works. The optional one, behind an EVIDENCE_BACKEND=gbrain flag, is a hybrid vector-and-keyword semantic store I ported from an earlier Dataclaw project, and it catches the case plain keyword search misses: "does linoleic acid raise cancer risk" and "are seed oils carcinogenic" are the same question wearing different words. Because both backends satisfy the identical interface, moving from keyword to semantic retrieval is a config change, not a rewrite.
Retrieval only finds candidates; deciding what each one is worth is a separate job, and that part is fully coded rather than aspirational. A claim taxonomy (causal-risk, benefit, superiority, mechanism, absolutist) sets what a good answer to that kind of claim should even look like. A study-design hierarchy (the evidence pyramid) assigns each design a quality value, from a meta-analysis at 1.00 and an RCT at 0.85 down through cohort at 0.60 to animal at 0.20 and in-vitro at 0.10. The ranker folds that together with human relevance, recency, sample size, directness and risk of bias into one transparent 0..1 score per study, and that score becomes the High, Medium or Low weight you see on each card. Every number in the report walks back to it.
I’ll be straight about what’s built and what isn’t. What runs today is a typed store with pluggable keyword or semantic retrieval and a ranking taxonomy that’s real and working. Where it’s headed is ontology-guided GraphRAG: storing each study as an exposure acting on an outcome, so retrieval can walk from a claim to evidence that never shared its vocabulary. I drew the interface the way I did specifically so that becomes the next backend, not the next rewrite.
| Layer | Job | Status |
|---|---|---|
| Typed store | Each study a structured Source: design, human relevance, directness, direction, weight, limitations | Built |
| Retrieval | SQLite FTS5 / BM25 by default; optional gbrain hybrid vector + keyword, same interface | Built, pluggable |
| Ranking taxonomy | Claim types, evidence pyramid, transparent 0..1 per-study score | Built |
| Ontology graph | Exposure–outcome graph traversal for vocabulary-independent retrieval | Next backend |
The pipeline, and what I don’t let the model do
A claim runs through a chain of deliberately narrow steps: parse it, decompose it, build the queries, retrieve from the store, classify each study, rank it, synthesize the direction, map that to a verdict, and run the safety layer. But the decision I care about most isn’t any one of those steps: it’s what I refuse to let the language model touch.
flowchart TD
A["User claim<br/>web or Telegram"] --> B["Claim parser"]
B --> C{"One claim<br/>or many?"}
C -->|multi-claim post| D["Decomposer<br/>split into subclaims"]
C -->|single| E["Normalize claim"]
D --> E
E --> F["Query builder<br/>generate PubMed queries"]
F --> G["Evidence store<br/>seeded + live PubMed if thin"]
G --> H["Study classifier<br/>type / human vs animal / directness"]
H --> I["Evidence ranker<br/>weighted score per study"]
I --> J["Evidence synthesizer<br/>direction + contradictions"]
J --> K["Verdict engine<br/>DETERMINISTIC over structured evidence"]
K --> L{"Safety gate<br/>high-risk topic?"}
L -->|yes| M["Escalate caveats + clinician referral"]
L -->|no| N["Standard disclaimer"]
M --> O["Report generator"]
N --> O
O --> P["Web report"]
O --> Q["Telegram short verdict + link"]
The model never grades evidence from memory, and I keep it boxed into three jobs, all of which I can audit. It parses messy input into testable claims. It labels each retrieved study’s stance (supports, contradicts, mixed, or not relevant), judging only from that study’s own finding text. And it writes the plain-English summary, but only after the verdict is already fixed. The verdict itself is a pure function over the direction tallies, the evidence quality, the directness and the claim type. Which means swapping the model can only ever move a verdict by relabelling an individual study, something you can check against the abstract yourself, and it can never cite a paper that wasn’t actually retrieved. That, more than anything else, is how I keep it from hallucinating a confident answer.
How a study earns its weight
| Scoring factor | Weight |
|---|---|
| Study design quality | 30% |
| Human relevance | 20% |
| Consistency across studies | 15% |
| Recency | 10% |
| Sample size / power | 10% |
| Directness to the claim | 10% |
| Risk of bias / conflicts | 5% |
The verdict engine reads and writes one structured object per claim, and it’s worth looking at the last field. "What would change the answer" is stored right alongside the verdict, generated as part of the same pass rather than tacked on, so the report always carries its own falsification conditions.
{
"primary_claim": "Vegetable oils cause cancer",
"verdict": "unsupported_as_broad_claim",
"confidence": "moderate",
"evidence_quality": "moderate",
"directness": "medium",
"evidence_breakdown": { "systematic_reviews": 4, "rcts": 1,
"cohort": 12, "animal_or_mechanistic": 18 },
"evidence_direction": { "supports": 2, "mixed": 7,
"contradicts": 10, "not_relevant": 4 },
"caveats": ["Repeatedly reused frying oil is a separate question",
"Oil types differ; the category is too broad"],
"what_would_change_the_answer": [
"Large prospective human studies showing dose-response harm",
"RCTs with disease-relevant outcomes by oil type"
]
}
The stack, one job each
OpenClaw runtime
The local-first, messaging-first agent runtime. It runs the Telegram gateway, conversational state, and the agent loop, self-hosted on hardware I control.
Pluggable evidence store
SQLite FTS5 (BM25) by default, optional gbrain hybrid vector + keyword, behind one interface. Each study stored as a typed object, not a blob.
Python + FastAPI core
The evidence brain: parse, retrieve, classify, rank, verdict, safety, report. Every decision lives here; the surfaces are thin clients.
Next.js + TypeScript web
The home page, report, and library, calling the core over HTTP. The Telegram bridge is a separate TypeScript OpenClaw plugin.
Seeded + live PubMed
~25 common claims seeded via the NCBI E-utilities ingest CLI; novel claims with thin coverage retrieve live from PubMed at query time and cache back.
Provider-agnostic LLM layer
Parsing and summary sit behind one interface, so the model vendor can change without touching the pipeline.
Deterministic verdict engine
Plain scoring code over the structured fields. No model call, and no OpenClaw agent, decides a verdict, which is what keeps citations honest.
The bits that fought back
The first thing I got wrong was assuming quality and directness were the same axis. A rigorous meta-analysis on vitamin retention in pressure-cooked vegetables is genuinely high quality, and it still barely answers "is pressure cooking bad for you." Until I split those into two separate scores, both computed and both shown, the report kept quietly overweighting a precise answer to the wrong question.
Then there was mechanism inflation, which turned out to be the whole game. Almost every viral health claim points at a real mechanism (oxidation, a hormone marker, an inflammatory pathway), and the mechanism usually is real. The hard part was making "this mechanism exists in a dish" and "it matters at the amount you eat at dinner" land on different verdicts, so a rat study about overheated oil never gets to read as proof your cooking causes cancer.
Pulling claims apart was harder than it looked too. Influencer posts weld four claims into one breathless sentence, and the extractor had to separate them without inventing a claim that wasn’t there and without dropping one that was. That’s a much narrower target than "just split on the conjunctions," and it took real iteration to hit.
I was also surprised by how much work the tone took. Writing verdict language that informs without frightening ate as many revisions as the ranking did. "Overstated" has to read as a fair, slightly boring assessment, not a scold, because the moment it feels like a scold, people stop trusting the tool on exactly the topics that scared them enough to check.
And then the trade I made on the database itself. Seeding a snapshot kept the demo deterministic and honest, with no hallucinated citations, at the cost of breadth. I made that call on purpose for a portfolio-stage build, and I made it survivable: because the store sits behind one interface, the default SQLite/BM25 backend, the optional gbrain semantic one, and live PubMed at query time are all the same shape to the pipeline. That’s why live retrieval already backfills a novel claim when the local shelf is thin, without anything downstream needing to change.
What it can and cannot do
Can, today
- Check nutrition, fitness, supplement, and sleep claims end to end.
- Decompose a bundled multi-claim post into separate checkable subclaims.
- Retrieve live from PubMed when a novel claim has thin local coverage, and cache it back.
- Show its evidence, its direction, and its uncertainty, never a bare label.
- Answer from Telegram and link to the full web report; save reports to revisit.
- Run its own eval harness: verdict-in-band accuracy and high-risk recall against a gold set.
Cannot, on purpose or not yet
- Cover PubMed broadly yet; the seeded store is ~25 common claims, with live retrieval only on the edges.
- Diagnose, or advise on medication, pregnancy, pediatric, or eating-disorder topics.
- Read screenshots, URLs, or video transcripts.
- Personalize to your labs, history, or medical conditions.
- Give a plain true or false; that is a design choice, not a gap.
Where I make it step back on purpose
A detector watches for the topics where a breezy evidence summary could actually do harm (pregnancy, children, cancer, kidney and liver disease, medication interactions, eating disorders, severe symptoms), and when it sees one it escalates the caveats or refuses to give treatment advice at all. Sift will never tell someone to stop a medication or hand out a personalized plan, because it has no tool that could. And every report carries the same disclaimer whether the claim was about frying oil or something far more serious, so the line never blurs.
I never wanted to tell anyone what to eat. I wanted to make "what does the evidence actually say" a one-minute question instead of an afternoon in PubMed.





