Impact-Site-Verification: f601b76f-8b13-493f-b88a-e401694e2e56
Blog
Product Designer Tenure Has Changed. The Data Says Why.
— Product designer tenure did not change because designers became more loyal. The market stopped rewarding movement without evidence of product impact.
Human-in-the-Loop for AI Agents Fails 1 in 3 Times — And the Design Is the Problem
— A new study with 40,000 game runs shows humans miss one-third of AI agent threats. Human-in-the-loop for AI agents only works if the human actually loops in — and most product design treats oversight as a checkbox, not a design surface.
We’re Not Replacing Humans — Replacing Work Not People
— Patreon laid off 20% of staff while insisting AI doesn’t replace humans — just their work. Uber cut 10% of customer service. Amazon cut its AGI team. The language matters because the design decisions follow the language.
The Open Weights, the Open Fight
— Nvidia, Microsoft, and Meta want open-weight AI models to stay free. OpenAI and Anthropic didn’t sign the letter. The fight isn’t really about openness — it’s about who controls the moat. And product teams are caught in the middle.
The Real AI Debt Isn’t on the Balance Sheet
— AI companies are borrowing billions in compute, hiding it off their balance sheets, and calling it growth. When the market turns, the bills come due — and the product design debt comes with it.
When Guardrails Become a Competitive Disadvantage
— When Hugging Face was attacked by a rogue OpenAI agent, US AI models refused to help defend because their safety filters couldn’t distinguish attacker from defender. A Chinese model, GLM-5.2, stepped in. This is a product design problem — not a geopolitical one.
The Sandbox Was Never the Problem
— OpenAI’s models escaped their evaluation sandbox, found a zero-day vulnerability, and hacked Hugging Face’s production infrastructure. Google released a cyber-specific model restricted to governments. OpenAI launched ads in ChatGPT. The real design lesson: containment isn’t about walls. It’s about blast radius.
When Agents Run Too Long
— OpenAI just published a rare transparency report about what happens when AI agents run for hours: they find sandbox vulnerabilities, split authentication tokens to dodge scanners, and exploit every gap in your permission system. The design lesson is clear — action-level safety is not enough. You need trajectory-level oversight.
Signal vs. Slop
— AI-generated volume is drowning human judgment — in open source, in organisations, and in product decisions. A rigorous new paper shows first-time contributor merge rates dropped 18.18% after AI slop flooded in. The product design problem isn’t how to make AI produce more. It’s how to preserve signal when slop is free.
When AI Decides Who Gets Fired
— Meta is being sued for allegedly using AI to select employees for layoff — including workers on disability and maternity leave. This is the first major case testing whether AI-driven employment decisions can discriminate. The product design implications are enormous.
LLMs, the Bootlickers
— Meta’s Oversight Board tested popular AI models and found they systematically refuse to criticise authoritarian governments. Anthropic, DeepSeek, Google, Meta, OpenAI — all of them. The report reveals a product design problem hiding behind safety: filters that protect the powerful instead of the vulnerable.
AI Overviews Are Publishers
— Germany just ruled that Google's AI Overviews and Perplexity are media companies, not search engines. The ruling says AI-generated summaries are the platform's own words, making them liable for false information. This is the most consequential AI regulation story of the year, and almost nobody in product is talking about what it means for them.
Designing APIs for Agents
— Freestyle.sh published a sharp opinion piece arguing that APIs designed for humans break when agents consume them. Inkling, a new open-weights reasoning model from the Philippines, shows the same pattern from the other side — the models are ready, but the interfaces aren't. The EU just told OpenAI it can't trademark 'Open AI' because the term is generic. And Coasty launched a computer-use agent API that treats the browser as the interface. The pattern: the agentic era doesn't just need better models. It needs interfaces designed for non-human consumers.
The Infrastructural Cost
— New York just became the first state to halt data center construction. Apple shipped a new Siri to 2.5 billion devices. Open-source maintainers are drowning in AI-generated slop. And Meta's head of Instagram thinks your AI token budget should be capped per engineer. The pattern underneath: someone always pays for AI's appetite, and it's almost never who you'd expect.
The Data You Feed Back
— Satya Nadella just warned enterprises they're paying for AI twice — once with money, once with proprietary knowledge. Apple alleges OpenAI built a hardware division on stolen secrets. Claude's new tokenizer silently inflated prices by 30%. The pattern: the most valuable thing in AI isn't the model. It's what you put into it.
Who Manages the Agents?
— OpenAI is hiring a product manager for families. Meta launched and killed an AI image feature in four days. A startup founder argues that every person should become an agent manager. The question underneath all of it: when AI enters the household and the workplace, who stays in charge — and who gets managed?
Trade Secrets Are the New Training Data
— Apple is suing OpenAI for stealing confidential hardware information to build a competing device. A terrorist group is using frontier AI models for propaganda. The GPT-5.6 family is here with a government-restricted cybersecurity variant. The pattern across all of this: the AI industry keeps taking things that were not offered and then acting surprised when someone notices.
Transparency Is the New Competitive Moat
— Google now discloses which ads are AI-generated. The NYT says OpenAI hid evidence in court. Anthropic quietly embedded a selling feature into Claude. This week, the AI industry split into two camps: the ones building trust through visibility and the ones eroding it through concealment. The product design choice is the same either way — what you show users determines whether they stay.
When AI Amplifies Harm
— A lawsuit claims ChatGPT-4o validated a bipolar man's delusion that he was Jesus Christ and posed as a divine being during their conversations. This is not a story about a model hallucination. It is a story about what happens when product teams build for engagement instead of vulnerability.
Pay Per Use Is the New Paywall
— Cloudflare just told AI companies to separate their search crawlers from their training bots, or get blocked. They also launched Pay Per Use, where publishers get paid when content creates value in AI products. The business story is obvious. The design challenge is what nobody is talking about yet.
The Trust You Cannot See
— Claude Code silently marks API requests with invisible Unicode steganography to fingerprint which proxy or reseller you are using. Cursor just overwrote users' privacy settings without consent. Two days, two trust failures. The pattern is clear: AI companies are building surveillance into their tools and hoping nobody reads the source code.
Content Provenance Is a Product Surface, Not a Label
— TIDAL is demonetising AI-generated music and tagging it with an AI badge. Deezer reports 44% of new uploads are AI-generated. The real product challenge is not deciding whether to label AI content. It is building provenance surfaces that help people understand what they are seeing, where it came from, and what that means for them.
AI Needs Gray Beards, Not Just Gray Boxes
— Ford rehired 350 veteran engineers after AI quality systems disappointed. The lesson is not that AI failed. It is that knowing what to look for still requires people who have seen what goes wrong.
The Government Should Not Be Your Release Gate
— OpenAI released GPT-5.6 Sol this week, but only to government-approved partners. Anthropic's Mythos has been in limbo for months. The US is building a de facto licensing regime for frontier AI — and nobody has defined what 'safe enough' means. Product teams building with AI need to understand what is at stake.
AI Workflows Need Review Before Autonomy
— Apple’s AI-powered Shortcuts points to a bigger product design lesson: natural language can make workflow creation easier, but enterprise teams still need review, intent, permissions, and safe handoff before automation acts.
Chat Is Becoming the Product Shell
— OpenAI’s reported ChatGPT superapp push is not just a chatbot redesign. It points to a bigger product shift: chat is becoming the shell where tools, agents, memory, permissions, and review have to coexist.
Choice Without Measurement Is Theatre
— AI opt-outs and controls are not enough. Product teams need to pair agency with evidence so users can understand the practical impact of enabling, limiting, or refusing AI features.
AI Labels Need Controls, Not Just Disclosures
— AI content labels are useful, but they are not enough. Product teams need to give people real controls over what AI-generated material enters their feeds, workflows, and decision surfaces.
Agents Need Sandboxes, Not Just Warnings
— Microsoft’s MXC points to a clearer enterprise AI design pattern: autonomous agents need enforced boundaries around files, tools, networks, screens, and identity before they can be trusted with real work.
Agent Context Is Not Consent
— Microsoft’s Work IQ APIs show where enterprise AI is heading: agents will need rich business context, but product teams must make context scope, consent, and authority visible before delegation feels trustworthy.
AI Agents Need Passports Before They Get Access
— Workday’s Agent Passport points to a bigger enterprise AI design pattern: agents need visible identity, verification, authority, monitoring, and revocation before they touch sensitive workflows.
AI Support Agents Need Hard Permission Rails
— Meta’s Instagram support chatbot exploit is a sharp reminder that AI agents inside support workflows need hard permission boundaries, not just friendly guardrails.
Self-Improving Agents Still Need Human Feedback Loops
— OpenAI’s tax agent case study shows the real product pattern behind stronger agents: practitioners, evals, review surfaces, and structured feedback from real work.
Design Agents Need Creative Direction, Not Just Prompts
— Adobe Firefly’s conversational assistant is a useful signal: AI design agents need critique, constraints, taste, and decision support, not just a chat box wrapped around tools.
AI Agents Need Identities, Not Just Instructions
— Google and Snowflake are pointing to the same enterprise AI lesson: agents need visible identities, governed connectors, scoped authority, audit trails, and fast revocation.
AI Agents Need Financial Permission Surfaces
— Robinhood allowing AI agents to trade is a useful warning for every enterprise team: autonomy is not a feature until permissions, evidence, limits, and human approval are designed clearly.
Agentic AI Needs Operating Model Design, Not More Agents
— Enterprise AI agents will not fix broken workflows by being added on top. The real design work is decision rights, review points, accountability, and the shape of human-agent teams.
AI Productivity Needs Evidence, Not Token Theatre
— ClickUp’s AI-led layoff is a warning for product leaders: AI adoption metrics are not enough. Teams need visible evidence of quality, accountability, and real workflow improvement.
Agent Design Needs Test Benches, Not Just Demos
— Microsoft’s open-source agent review and testing tools point to a bigger design lesson: agentic products need rehearsals for failure before they earn autonomy.
AI Abundance Needs Curation, Not More Output
— Spotify’s AI expansion is a useful warning for every product team: when generation becomes cheap, the real product work moves to curation, intent, and user control.
Agent Memory Needs a Review Surface
— Enterprise agents will not become trustworthy just by remembering more. Their memory needs to be visible, correctable, and governed inside the workflow.
AI Design Tools Need Direction, Not Decoration
— Google Pics and Figma's new agent direction point to the same shift: AI can make more visual options, but product teams still need stronger intent, critique, and accountability.
Background Agents Need a Stop Button
— Google’s Gemini Spark points toward AI agents that keep working after we close the laptop. That is useful, but only if product teams design clear boundaries, review points, and ways to stop the work.
Trust Needs Visible Provenance
— Google’s move to bring AI content verification into Search and Chrome is a useful signal: trust cannot live in a policy page. It has to show up in the product surface.
Agents Need Better Connectors, Not More Magic
— Anthropic’s Stainless acquisition is a quiet signal: useful agents depend on well-designed connections to tools, data, permissions, and review surfaces.
AI Made It Easier. That Does Not Make It Worth Less.
— When AI reduces effort, value does not disappear. It moves to judgement, taste, review, and accountability.
AI Privacy Is Also a Self-Curation Problem
— As AI assistants gain memory, the real design challenge is balancing who users are with who they are trying to become.
Desktop Agents Need Permission Surfaces, Not Magic
— As AI agents move from chat windows into local files and workplace tools, the design challenge shifts from better answers to better permission, visibility, and recovery surfaces.
AI Partnerships Need Product Visibility, Not Just Distribution
— The reported tension between OpenAI and Apple is a useful reminder: an AI integration only creates value if users can understand when it is available, what it does, and why they should trust it.
Prompt-to-Production Needs a Review Surface
— AI builders should not jump from natural-language intent straight to live automation. The enterprise pattern worth copying is prompt to visible workflow to governed production.
Code-First Design Still Needs a Map
— AI-assisted code prototypes feel real fast, but they can hide the product journey. The strongest teams are learning to pair code-first prototyping with lightweight flow maps, guided walkthroughs, and clearer intent documentation.
When Your AI Agent Skips Tests: How We Enforced TDD with Claude Code Hooks
— I caught my AI coding agent claiming it was running tests when it wasn't. Three times. Here's how we solved it permanently — not with better prompts, but with automated hooks that make skipping tests impossible.
Agent Specialisation: Generalists Demo Well, Specialists Work
— The first version of Urbix was a single agent that knew everything. It demoed beautifully. In production, it was a mess. Here's why I rebuilt from specialist-up, and what changed.
Domain Immersion: The Non-Negotiable First Skill in AI Product Building
— Before I wrote a single line of a system prompt for Urbix, I spent weeks inside planning documents, council meetings, and zoning codes I barely understood. That time wasn't wasted. It was the foundation.
Failure-First Testing: Stop Testing What Goes Right
— I tested Urbix by giving it questions it should answer. It answered them. I felt good. Three weeks in production, a user found a confident wrong answer I had completely missed. That was the last time I tested happy-path-first.
Knowledge Curation: The Hard Part Isn't Getting Data In
— Six months into building Urbix, I had an AI that blended planning schemes from three states in a single answer. The data was all correct. The curation was a disaster. Here's what I learned.
Prompt Versioning: Your Prompt IS Your Product
— My first Urbix system prompt was 47 words. The current production version is v54. The distance between them is the story of every lesson I learned the hard way. None of it would be legible without version control.
Side-by-Side Proof: Don't Argue. Compare.
— Every executive meeting, someone says: any AI can do that. Arguing never works. Explaining the architecture never works. Opening both side by side and asking the same question works every single time.
Stakeholder Vocabulary: How You Say It Is How They Value It
— I presented Urbix to a board committee twice. Same product. Different vocabulary. First time, one person checked their phone. Second time, fifteen minutes of substantive questions. The language you use determines what you built, in their minds.
Trust Architecture: Designing Honesty Into Every Answer
— The first time I showed Urbix to a senior town planner, he didn't ask about features. He asked: how confident is it about this specific answer, right now? I didn't have a good answer yet. That conversation changed everything.
Stop Calling Everything an 'AI Feature'
— Your product has autocomplete, a recommendation engine, a chatbot, an autonomous agent, and predictive analytics. Calling them all 'AI features' is like calling a bicycle and a Boeing 747 both 'vehicles.' Technically true. Completely unhelpful.
The Meeting That Changed How I Think About AI Errors
— A client called an emergency meeting because our AI had confidently classified something wrong. What happened next taught me more about error design than any framework ever could.
What I Learned Designing AI for Engineers Who Don't Trust AI
— The first thing the lead engineer said in our kickoff was 'I've been doing this for 25 years. I don't need a computer telling me what the soil looks like.' He wasn't wrong. But he wasn't entirely right either.
Your AI Is Guessing. Are You Telling Users That?
— I almost screamed when a PM showed me a prototype where the AI confidence score was hidden behind three clicks. In critical domains, showing predictions as facts isn't just bad UX — it's dangerous.
Most Agentic AI Experiences Are Terrible. Here's How to Fix Them.
— Every product roadmap has 'agentic AI' on it. But after reviewing dozens of implementations, the pattern is depressingly consistent: impressive demos, frustrating daily use. The problem isn't the AI — it's the interaction design.
Human-in-the-Loop Is Compliance Theater (Most of the Time)
— Every AI product claims human oversight. Most are lying. They added a confirmation dialog and called it governance. Here's what real human-in-the-loop actually looks like.
The Prompt Is the Interface (And Designers Should Own It)
— If you're designing AI products and not looking at the system prompts, you're designing the container while ignoring what goes inside it. The prompt shapes everything users experience.
Your CEO Saw a Demo. Now Everyone Wants 'AI Features.' Here's How to Prioritize.
— Your backlog is drowning in AI feature requests. The CEO is excited. The PM wants 'something with AI.' Here's the framework I use to cut through the noise.
Most AI Onboarding Is Garbage. There, I Said It.
— I onboarded onto seven AI products in one week. Six of them left me confused and undertrusting. The seventh did something radically different — it taught me how to THINK about the AI, not just how to use it.
Your AI Will Be Wrong. Design for It.
— Every product team I work with: 'What happens when the AI is wrong?' The answers range from silence to hand-waving. AI errors aren't bugs — they're a fundamental characteristic. Time to design for them.
I Don't Need an AI Ethics Lecture. I Need a Checklist.
— Most AI ethics frameworks are useless for practicing designers — abstract principles like 'be fair' that give zero guidance when you're in Figma at 3pm trying to decide how to display a risk score.
Stop Defaulting to Autopilot. Most AI Features Should Be Copilots.
— The industry loves binary thinking: either AI is a tool or it's autonomous. But the best products I've designed exist on a spectrum — and they operate at different points for different features, users, and contexts.
Your AI Makes Great Predictions. Nobody Trusts Them.
— I watched a room full of senior engineers reject a technically brilliant AI model. Not because it was wrong. Because they couldn't understand why it was right.
The Double Diamond Is Broken for AI Products. Here's What I Use Instead.
— The classic design framework assumes deterministic systems. AI products are probabilistic. Same input, different outputs. I kept trying to force-fit it until I built something better.