Case Study 3 of 4SHIPPEDEcolab Virtual Agent Design System & Obliq

An AI-Native Design Process

In 2025, engineering could build a working screen with AI in an afternoon while design still answered in two weeks, and Ecolab’s design system offered designers no guidance at all.

  • I designed an AI-native process for Ecolab’s 80 to 100 in-house designers.
  • I proved it by generating a design system with Claude connected to Figma in 60 to 80 hours, a person approving every change.
  • The process shipped.

Explore the Obliq Design System in Figma (opens in a new tab)

The cover page of the Obliq Virtual Agent design system: the Obliq asterisk logo and title, version 0.1.0 as a proof-of-concept build from June to July 2026 with 60-plus pages, tags for Tokens and Variables, Light Mode, MUI Ready, WCAG 2.2 AA and Draft, and a contents list of seven sections from Introduction to Onboarding Patterns.

At a glance

  • 80–100

    In-house designers who would use the AI-native design process I made the case for, built to speed up design-system maturity, product maturity and time to launch

  • ~2 wk

    To C-suite consensus on a new AI brand and identity that had stalled, which unlocked its rollout across the company

  • 60–80 h

    Of my own time to the first version of Obliq I showed, a generated design system with tokens, variables and guidance inside the file; 90 to 120 hours by the second round

The Decision

What I Chose

I made the case for an AI-native design process for Ecolab’s 80 to 100 in-house designers, to:

  • Speed up design-system maturity, product maturity and time to launch
  • Make feasibility clearer earlier
  • Make collaboration with outside agencies easier

My proof was a design system generated by connecting Claude and Figma through an MCP server (the connection that lets a model reach a tool or a data source) that a company builds and governs itself. Retrieval grounds the model in the tokens, components and rules that already exist, and a person approves every change before it is committed. I built it on my own machine first, so the case could be made with a working system instead of a slide.

My reasoning, in the order I would give it to a CTO:

  • Nothing new to buy: Claude and Figma were already in Ecolab’s stack.
  • The company stays in control: a self-built server means it sets the security model, sees how its data is handled and keeps its own audit trail.

What I Rejected, and Why

  • Hiring more designers. It is linear and slow, and it never changes the cycle: engineering would still build in an afternoon what design answers in two weeks.
  • Buying a design-to-code service. Fast to start, but opaque about how the company’s data is handled, with no audit trail the company controls, on someone else’s roadmap.
  • Neither was an AI strategy decision; both were ways to avoid making one. The path I chose had its own cost, covered at the end, and I would make the same call.

Context & Constraints

Three pages of Obliq's human-in-the-loop patterns library: starting with a real decision and choosing the right kind of human involvement; setting scope before the agent acts, with bounded delegation, a clarified goal and a permission boundary; and making execution visible and interruptible, with a live execution ledger and pause and stop controls.

Constraints

  • No Designer Guidance Since 2018
  • Two-Week Design Cycle
  • 8–10-Week Platform Launches
  • Neptune Team of 3–4
  • Own Time, Own Machine

The assistant in Case Study 1 needed a brand and a design system underneath it, and Ecolab didn’t have one built for AI. This case study is how I built that layer by hand, then proved a mature system could be generated with AI.

The gap was speed.

  • Engineering had Claude Code and could build a working screen in an afternoon; design had Figma and a two-week cycle, and products stalled at that seam.
  • Whole AI platforms shipped in eight to ten weeks with no AI design process built for that speed.
  • Teams worked in silos and fell back on familiar patterns, because nobody could tell them fast enough what was feasible. I was the bridge between design and engineering, and one bridge doesn’t scale.

The starting point made it harder. Neptune, Ecolab’s core design system since 2018, never reached what the industry would call maturity:

  • Very few variables, no tokenization and no accessibility standards
  • Components without auto-layout, hard to edit or extend
  • Every renaming and new component done by hand, and decisions made in meetings never logged
  • No onboarding or designer guidance, because the philosophy was that a design system only needs to guide developers

Its team of three was taking on debt faster than it could pay it down.

Guidance Is What Makes It a System

I built the Ecolab Virtual Agent layer on top of Neptune by hand and guidance first, and I worked the AI brand out in the room, because both needed people to agree.

Two pages of Obliq's guidance on notifications, alert rules and escalation: a table of who controls what, a My notifications settings screen, an optional SMS activation step, and an admin screen that defines when an alert should exist, with a rule preview and an AI-suggested rule change that a person reviews.

The Ecolab Virtual Agent AI & Agentic Design System was a layer on top of Neptune, and the first of its kind at Ecolab:

  • 68 pages, built by hand
  • No tokens or variables of its own; its components were the logo in its lockups
  • Around them, what designers had been asking for all along: context, resources, education and guidance, including the brand guidelines and the company’s first in-product brand guidelines, for a sister brand designed to live inside Neptune
  • Every good and bad example taken from real designs by the teams closest to AI product work

A group of designers were delighted with the help it gave them, which showed that listening to designers is efficient, not indulgent.

Brand and identity are very human: they carry emotional states, or the story of moving through them on the way to a purchase. So I worked the brand out in the room, speaking as a user in each session and asking each person questions as if I were them. It is a little performative, but it gets people to empathize fast. The AI brand had been in progress a long time, with ongoing agency spend, and leadership wanted it resolved. Within about two weeks I had C-suite consensus, which unlocked its rollout across the company, and brand and product leadership considered the sister brand a success because it stayed consistent inside Neptune.

This was the part that had to be built by hand. The rest of this case study is about the parts that didn’t.

An LLM Should Never Do a Job an If-Statement Can Do

I designed the pipeline so the model proposes and a person decides: propose, ground, gate, commit, with a blocked state that explains itself and a drift check that flags but never fixes.

Scenario
  1. Propose

    Claude drafts the token, style, or component.

  2. Ground

    RAG checks it against the tokens and rules that already exist.

  3. Gate

    A deterministic check. An if-statement, not a model.

  4. Approve

    One person reviews it and says yes.

  5. Commit

    Written to the Figma file.

  6. Log

    What changed, who approved it, and when.

Blocked, and It Shows Its Work

  • What was caught
  • how it was caught
  • why it matters
  • ways forward

Then back to propose, with reasons.

Drift Checker

Watches the file after commit and flags anything that drifts from the system. It never auto-fixes; the user decides.

Press Step to send a proposal through.

Judgment where it's needed. Certainty everywhere else.

  1. Propose

    Claude drafts the token, style, or component.

  2. Ground

    RAG checks it against the tokens and rules that already exist.

  3. Gate

    A deterministic check. An if-statement, not a model.

  4. Approve

    One person reviews it and says yes.

  5. Commit

    Written to the Figma file.

  6. Log

    What changed, who approved it, and when.

  • Blocked, and It Shows Its Work

    What was caught · how it was caught · why it matters · ways forward. Then back to propose, with reasons.

  • Drift Checker

    Watches the file after commit and flags anything that drifts from the system. It never auto-fixes; the user decides.

  • User
  • Agent
  • Gate
  • Log

Judgment where it's needed. Certainty everywhere else.

The diagram is the pipeline. What sits underneath it:

  • Grounded, not remembered. Retrieval (RAG) checks each proposal against what already exists, not the model’s memory of what a design system usually looks like. The gate then checks naming, scope and bindings.
  • Flag, never fix. The drift checker never fixes anything on its own, because silent fixes are how design systems rot without anyone noticing.
  • Ownership is split on purpose. Visual intent lives in Figma, token values live in the code repository and sync to Figma, component identity and policy live in a small registry, and the shipped appearance lives in code. The drift checker sits between Figma and the registry, with a person at the decision.
  • Blocked states show their work: what was caught, how, why it matters and the ways forward. It is the same pattern I would put in front of any user (AI proposes, a person decides), and here the user is the designer.

The operating layer is the least glamorous part and, I think, the most convincing: a CLAUDE.md instruction file and, inside the Figma file, a “working on this file” page with standing rules, page anatomy, an ID map, a definition of done and a log of the model’s mistakes, so every new session starts smarter than the last. This portfolio was built the same way.

Brand by Hand, System by Machine

I compared the two methods honestly: the hand-built system was better where people had to agree, and the generated one was better where rules had to hold.

The hand-built Ecolab Virtual Agent layer in Figma: a Brand Guidance page with logo guidelines for the full-length logo, the logomark, the short logo and the sparkle element, each with do and don't examples, beside Brand Components, Forms and Fields, and Cards pages.
ECOLAB (MANUAL)
Pages of the generated Obliq system in Figma: About Obliq Virtual Agent (what it is, three kinds of AI, why agentic AI is different, with a consent card example), the AI and Agentic Principles and Best Practices page, and an Accessibility page with WCAG 2.2 AA color pairings and token upgrades.
OBLIQ (MCP)

The two systems weren’t the same kind of thing, so the fair comparison is by strength, with every line a verified number or a plain fact.

Hand-built Ecolab Virtual Agent layer

  • More than a year of manual effort, spring 2025 to July 2026, alongside a full-time product role
  • 68 pages, no tokens or variables, and the logo lockups as its only components
  • Weaker than Neptune as a system, but better at brand nuance and needed fewer naming corrections, because the brand was worked out with people, not generated

Obliq

  • 60 to 80 hours (about two weeks) for the first version, 90 to 120 for the second, and a far more complete system
  • 95 pages, 13 text styles applied across the file and 8 main components after an audit
  • Variables bound by construction, a drift checker in the loop and guidance for designers and developers inside the file
  • Hand-off from Figma through MCP to React, with MUI themed by the same tokens: design owns the skin, engineering owns the skeleton
  • Components refined first and converted later; that work wasn’t finished

Where the User Stays in Charge

The same human-in-the-loop system that keeps a user in charge in Case Study 2 keeps a design system safe to generate. Here the user is the designer:

  • Disclosure of what the agent is about to do
  • A gate before anything is committed
  • A blocked state that explains itself
  • Feedback that changes the rules
  • A log a third party could audit

Evidence & Outcome

  • ~2 Weeks to C-Suite Consensus

    The AI brand had stalled until I took it into facilitated sessions. About two weeks later the C-suite agreed, which unlocked the rollout of the new AI brand and identity across the company, as a sister brand inside Neptune.

  • 60–80 Hours to the First Version

    I showed the first version of Obliq to design leadership, then presented it informally to the AI product owner and formally to product leadership, with engineering leadership in the room. It was the proof that a team could get a usable system in weeks, not years, without waiting for perfection first.

  • From Proof of Concept to Shipped

    When I told my AI product owner what I had been building, she named the missing piece in another team’s work: a prototyping agent (Claude connected to Figma through MCP) that Ecolab’s AI, data science, data analytics and engineering teams were building with Product, so product leaders could test an idea with one prompt before sharing it more widely. It needed a design system up front, and I had one. The hand-built Ecolab Virtual Agent layer connected it, the process I had presented as a proof of concept shipped as the real thing, and I trained the designers on it.

Hours are from my own logs. Page counts and the brand timeline are from my project records. Obliq counts are pulled from the Obliq file. Obliq is my own IP and was never part of Ecolab’s architecture; the process that shipped ran on the hand-built Ecolab Virtual Agent layer.

What It Cost, and What I’d Do Differently

What It Cost

  • The hours were real, and so were the model’s mistakes, enough of them to justify a mistakes log.
  • The coded path (React with a themed MUI library) runs as a prototype, not a shipped product.
  • I built Obliq on my own machine, on my own time, to make a case the roadmap hadn’t made room for. That is proof of conviction, and also a sign of a process that shouldn’t have needed it.

What I’d Do Differently

  • Start the charter and the mistakes log on day one, not after the first bad session, because the rules a model follows are a design deliverable and deserve to be treated like one.
  • Define the component prop contracts before generating, so the gate has something exact to check.
  • Measure rework hours, not just build hours. That is the number that would have made the business case impossible to argue with, and I didn’t track it.
  • Find out what other teams are already building before I start, because I built in parallel with a team that needed exactly what I had.