Software Development

Give Claude Eyes: A Screenshot Feedback Loop for Better AI UI

Tech Arion TeamTech Arion Team
October 1, 202614 min read0 views
Give Claude Eyes: A Screenshot Feedback Loop for Better AI UI
The biggest quality jump in AI-built UI comes from letting the model see what it made. Set up Playwright MCP or Chrome DevTools MCP, a three-width critique rubric and automated gates.

The single biggest quality jump you can give Claude on front-end work is letting it see what it built. Install Playwright MCP (claude mcp add playwright npx @playwright/mcp@latest) or Chrome DevTools MCP, have Claude screenshot the page at 375, 768 and 1440 pixels wide, critique each capture against a fixed rubric, fix what it finds, and repeat until the screenshots pass. Anthropic's own Claude Code best practices put it bluntly: give Claude a check it can run, such as a screenshot to compare, because without one 'looks done' is the only signal it has. This guide walks through the tooling, the exact loop we use, a copy-paste critique rubric and the automated gates that catch what eyes miss.

Why does Claude ship UI that looks broken when it says it is done?

A language model writing JSX is working blind. It predicts markup and class names, but it never sees the rendered result: the hero headline that wraps into four lines on a phone, the card grid that leaves an orphan, the grey-on-grey caption that fails contrast, the sticky header that hides the first heading. The Chrome team framed the problem exactly when it launched Chrome DevTools MCP: coding agents 'are not able to see what the code they generate actually does when it runs in the browser.' Anthropic's best practices page says Claude stops when the work looks done, and without a runnable check you become the verification loop, noticing every mistake yourself. Prompts and design rules (see our pillar post Why AI Websites Look the Same) set the direction; screenshots are how the model finds out whether it got there.

  • •Text-only review catches syntax and logic, not overflow, collisions, cramped spacing or weak hierarchy.
  • •Most layout bugs only appear at one breakpoint, so a single desktop check misses them.
  • •A screenshot turns 'make it look better' into a concrete diff the model can list and fix.
  • •Anthropic recommends asking Claude to show evidence, such as a screenshot, rather than asserting success.

What does Anthropic recommend for verifying UI changes visually?

The Claude Code best practices documentation lists 'Verify UI changes visually' as a core strategy. Its before example is the vague 'make the dashboard look better'. Its after example is: paste a screenshot, 'implement this design. take a screenshot of the result and compare it to the original. list differences and fix them.' The page names a browser screenshot compared against a design as one of the checks that closes the loop, alongside tests, build exit codes and linters, and suggests three ways to make the check stick: ask for it in the same prompt, set it as a /goal condition that a separate evaluator re-checks after every turn, or enforce it with a Stop hook that blocks the turn from ending until a script passes. It also recommends a second opinion: a verification subagent in a fresh context grades the work, so the agent that wrote the code is not the one marking it.

Playwright MCP vs Chrome DevTools MCP vs Claude in Chrome: which should you use?

Three tools give Claude a browser today, and they overlap more than they differ. Playwright MCP, from Microsoft, is the most common choice for local development and CI-like runs. Chrome DevTools MCP, from the Chrome team, adds DevTools depth: performance traces, CSS inspection and a Lighthouse audit tool. Claude in Chrome is Anthropic's browser extension, connected to Claude Code with the --chrome flag, and it drives your real, signed-in Chrome window. Microsoft also now points coding agents to playwright-cli, a command-line version exposed as skills, because CLI calls avoid loading large tool schemas and accessibility trees into context. Pick one primary tool per project and write its name into your CLAUDE.md so every session uses the same one.

ToolHow you install or enable itScreenshot and resize toolsExtra strengthsBest for
Playwright MCP (Microsoft)claude mcp add playwright npx @playwright/mcp@latestbrowser_take_screenshot (fullPage option), browser_resize, browser_snapshotAccessibility-tree snapshots, --device and --viewport-size flags, headless or headed, optional --caps visionDay-to-day build, screenshot and fix loops on localhost
Playwright CLI + skills (Microsoft)npm install -g @playwright/cli@latest, then playwright-cli install --skillsplaywright-cli screenshot, --hires for full device pixel ratioMore token-efficient than MCP, per Microsoft's READMELong sessions where context is tight
Chrome DevTools MCP (Google)claude mcp add chrome-devtools --scope user npx chrome-devtools-mcp@latesttake_screenshot, take_snapshot, resize_page, emulatelighthouse_audit, performance_start_trace, get_css_styles, console and network inspectionLayout debugging, performance and accessibility audits
Claude in Chrome (Anthropic)Install the extension, run claude --chrome or enable it via /chromeScreenshots in your visible Chrome window, save to diskShares your login state, records GIFs, reads console errorsAuthenticated apps, staging behind login, design verification against a Figma mock

Snapshot or screenshot: which one actually helps design quality?

This trips up most people setting up Playwright MCP. By default it works from Playwright's accessibility tree, not pixels; its README says it needs no vision model and operates on structured data. The browser_snapshot tool returns that tree and is what Claude uses to click and type. The tool description for browser_take_screenshot is explicit: you cannot perform actions based on the screenshot, use browser_snapshot for actions. So snapshots are for navigation and semantics, and screenshots are for judging how the page looks. For design review you need the screenshot, because spacing rhythm, alignment and contrast do not exist in an accessibility tree. Use both: the snapshot reveals a button with no accessible name, the screenshot reveals it sits off the grid. Chrome DevTools MCP mirrors the split with take_snapshot and take_screenshot. Coordinate-based clicking in Playwright MCP is an opt-in extra under --caps vision; you do not need it for visual review.

  • •browser_snapshot: structure, roles, labels and element refs. Cheap, precise, blind to style.
  • •browser_take_screenshot with fullPage true: the whole scrollable page, for hierarchy and rhythm.
  • •scale 'css' keeps images small and consistent; 'device' gives high-resolution captures when you need to inspect type rendering.

How do you set up the screenshot, critique, fix loop step by step?

Here is the loop we run on landing pages and marketing sites, written so you can paste it into a project. It assumes a dev server on localhost:3000 and Playwright MCP, but the same steps work with Chrome DevTools MCP by swapping tool names (resize_page and take_screenshot). Keep the loop bounded, around three rounds, so old screenshots do not crowd the context; after two failed corrections on the same issue, Anthropic's advice is to clear context and restart with a sharper prompt.

1
Install the browser tool once

Run claude mcp add playwright npx @playwright/mcp@latest, then /mcp in Claude Code to confirm it is connected. Add --headless to the args if you do not want a visible browser window, or --device "iPhone 15" for a dedicated mobile server.

2
Give Claude a visual target

Paste a reference screenshot, a Figma frame or a written art direction, plus your design rules file. Without a target, the model grades against its own defaults, which is exactly how you end up with the generic look. Our sibling posts on Claude UI design prompts and on design tokens in CLAUDE.md cover how to write this target.

3
Capture at three widths

Prompt: 'Start the dev server if needed. For each of 375x812, 768x1024 and 1440x900: browser_resize, navigate to http://localhost:3000/, wait for fonts and images, take a fullPage screenshot named home-<width>.png.' Three files, one per breakpoint, every round.

4
Critique against the rubric, not vibes

Prompt: 'Review each screenshot against the rubric in docs/design-review.md. For every failure, give the width, the element, what is wrong and the rule it breaks. Rank as Blocker, High, Medium or Nit. Do not edit code yet.' Separating critique from fixing stops the model rationalising its own layout.

5
Fix the top issues only

Prompt: 'Fix all Blockers and High items. Change tokens or shared components rather than adding one-off pixel values.' Small batches keep diffs reviewable.

6
Re-capture and compare

Repeat the capture at all three widths, then: 'Compare new screenshots with the previous round. List what changed, what is still failing, and anything that got worse.' Regressions at another breakpoint are common after a mobile fix.

7
Hand over evidence

Finish with the final screenshots, the console messages (browser_console_messages) and the list of remaining Medium and Nit items. You review images and a short list, not a wall of reassurance.

Why test at 375, 768 and 1440 pixels?

These three widths cover a modern phone, a portrait tablet and a typical laptop or desktop canvas, and they straddle the breakpoints most Tailwind and CSS frameworks use, so a layout that survives all three rarely breaks badly in between. They are also what Patrick Ellis's open-source design-review agent for Claude Code (the OneRedOak claude-code-workflows repository) tests: a 1440 pixel desktop viewport with a screenshot, 768 for tablet layout adaptation, and 375 for touch optimisation, with an explicit check for no horizontal scrolling or element overlap. Do not skip 768: tablet is where two-column grids squeeze and navigation menus collapse awkwardly. The point is a fixed list that runs every time, and in Playwright Test you can encode it as projects using device presets such as devices['iPhone 13'] or a custom viewport object.

What should Claude check in each screenshot? A design critique rubric

A rubric turns taste into checks a model can apply consistently. Save this as docs/design-review.md, reference it from CLAUDE.md, and have the critique step cite rule names. It borrows from Nielsen Norman Group's definition of visual hierarchy (organising elements so the eye consumes them in order of importance, using color and contrast, scale and grouping) and from WCAG 2.2 for the measurable parts.

Screenshot critique rubric

Hierarchy: within three seconds, is it obvious what the page is and what to do next? One dominant headline, one primary CTA per viewport, no more than about three type sizes competing in a section.
Spacing rhythm: gaps come from a fixed scale (for example 4, 8, 12, 16, 24, 32, 48). Related items sit closer together than unrelated ones. Section padding is consistent from top to bottom.
Alignment: edges line up on a shared grid. Text blocks, card contents and buttons share left edges. Nothing is centred by accident next to left-aligned siblings.
Contrast: body text reaches at least 4.5:1 against its background and large text at least 3:1 (WCAG 1.4.3). Check text over images and gradients specifically.
Overflow: no horizontal scroll at 375px, no text clipped by fixed heights, no long words or URLs breaking containers, no overlapping elements.
Tap targets: interactive targets are at least 24 by 24 CSS pixels or adequately spaced (WCAG 2.5.8), and primary actions on mobile are comfortably larger; 2.5.5 sets 44 by 44 as the enhanced level.
Consistency: buttons, cards and form fields reuse the same radius, border, shadow and colour tokens across the page.
Content realism: real-length copy, not lorem ipsum that hides wrapping problems.
Console: zero errors and no failed image or font requests on the captured page.

Which automated gates should back up the visual loop?

Model critique is good at hierarchy and rhythm and unreliable at exact measurements, so pair it with deterministic tools that return pass or fail. Each one below is something Claude can run and read in the conversation, which is the property Anthropic's guidance asks for. Wire them into a Stop hook or CI so a session cannot finish with them red.

  • •Playwright visual regression: await expect(page).toHaveScreenshot() saves a baseline on the first run and fails on pixel drift afterwards; update intentionally with npx playwright test --update-snapshots, and set tolerance with maxDiffPixels. Generate baselines in the same environment you test in, since Playwright warns rendering varies with OS, browser version, hardware and headless mode.
  • •Accessibility scans: @axe-core/playwright with new AxeBuilder({ page }).withTags(['wcag2a','wcag2aa','wcag21a','wcag21aa']).analyze(), asserting violations equal an empty array. Deque says axe-core finds on average 57 percent of WCAG issues automatically, and Playwright notes automated testing cannot detect every violation, so keep the human pass.
  • •Lighthouse: Chrome DevTools MCP's lighthouse_audit tool scores accessibility, SEO and best practices from inside the session (performance needs performance_start_trace). For CI, Lighthouse CI (npm install -g @lhci/cli, then lhci autorun) can fail builds when scores regress.
  • •Core Web Vitals budget: web.dev's good thresholds are LCP within 2.5 seconds, INP of 200 milliseconds or less and CLS of 0.1 or less, measured at the 75th percentile. Our post on Framer Motion and GSAP animations with Claude covers keeping CLS in check.
  • •Console: fail the run on any console error or 404 asset.

What goes wrong with AI visual testing, and how do you avoid it?

The loop is simple, but a few failure patterns show up again and again when teams first switch it on.

⚠️Only checking the desktop viewport

Consequence: The page looks polished at 1440px while the phone layout has horizontal scroll and a hidden CTA, and most traffic sees the broken version.

Solution: Hard-code the three widths into the prompt or CLAUDE.md so every round captures all of them.

⚠️Letting the same context write and grade the design

Consequence: The model rates its own layout generously and declares victory after cosmetic tweaks.

Solution: Run the critique in a fresh subagent with only the screenshots and the rubric, as Anthropic suggests for adversarial review.

⚠️Screenshotting before fonts, images and animations settle

Consequence: Fallback fonts and half-faded sections produce false findings, and visual regression baselines flake.

Solution: Wait for network idle and fonts, disable or finish entrance animations, and use a stylePath CSS file to hide volatile elements in toHaveScreenshot.

⚠️Relying on the model for contrast and target sizes

Consequence: Eyeballed estimates pass text that fails 4.5:1 or buttons below 24px.

Solution: Measure with axe-core or Lighthouse and let the model handle what tools cannot judge: hierarchy, rhythm and taste.

How do you make the loop automatic in every Claude Code session?

Instructions only help if they run every time. Put a short Visual verification block in CLAUDE.md: after any front-end change, capture the changed routes at 375, 768 and 1440, critique against docs/design-review.md, fix Blockers and High items, and attach final screenshots. The OneRedOak repository ships a similar CLAUDE.md snippet (navigate to each changed view, compare against design principles and a style guide, take a full-page desktop screenshot, check console messages) plus a design-review subagent that triages findings as Blocker, High-Priority, Medium-Priority or Nitpick. R.A. Parks describes the same pattern on his blog, with a reviewer subagent that produces graded reports. For hard enforcement, Anthropic documents Stop hooks that block a turn from ending until your script passes, which is the right home for toHaveScreenshot and axe checks. At Tech Arion we keep the rubric and the viewport list in the repo next to the design tokens, so every session, human or agent, reviews against the same rules; if you want that setup on your own site, our vibe coding service at techarion.com/services/vibe-coding includes it.

Frequently asked questions

Short answers to the questions developers ask most when wiring screenshots into Claude.

Frequently Asked Questions

Want AI-built pages that pass a real design review?

Tech Arion builds websites with Claude Code and a visual QA loop: design rules in the repo, screenshots at every breakpoint, accessibility and performance gates, and a human signing off before launch. Tell us what you are building and we will show you how the loop fits your stack.

Sources & References

Sources fetched on 1 October 2026:

  1. 1.

    Anthropic. Best practices for Claude Code - Claude Code documentation. 'Give Claude a way to verify its work', 'Verify UI changes visually', /goal, Stop hooks and verification subagents.

    View Source
  2. 2.

    Anthropic. Use Claude Code with Chrome - Claude Code documentation. --chrome flag, /chrome, extension 1.0.36+, design verification and screenshots.

    View Source
  3. 3.

    Microsoft. Playwright MCP - GitHub repository and README. Install command, accessibility snapshots, browser_snapshot, browser_take_screenshot, browser_resize, --caps vision, --device, --headless.

    View Source
  4. 4.

    Microsoft. Playwright CLI - GitHub repository. Token-efficient CLI with skills for coding agents.

    View Source
  5. 5.

    Bynens, M. and Hablich, M. (23 Sep 2025). Chrome DevTools (MCP) for your AI agent. Chrome for Developers blog.

    View Source
  6. 6.

    Google Chrome DevTools team. Chrome DevTools MCP tool reference: take_screenshot, take_snapshot, resize_page, emulate, lighthouse_audit, performance_start_trace.

    View Source
  7. 7.

    Playwright. Visual comparisons - toHaveScreenshot, --update-snapshots, maxDiffPixels, stylePath and rendering-variance warning.

    View Source
  8. 8.

    Playwright. Accessibility testing with @axe-core/playwright and AxeBuilder.

    View Source
  9. 9.

    Deque Systems. axe-core - accessibility engine; finds on average 57% of WCAG issues automatically.

    View Source
  10. 10.

    W3C WAI. Understanding SC 2.5.8 Target Size (Minimum), WCAG 2.2.

    View Source
  11. 11.

    W3C WAI. Understanding SC 1.4.3 Contrast (Minimum), WCAG 2.2.

    View Source
  12. 12.

    web.dev. Web Vitals - LCP 2.5 s, INP 200 ms, CLS 0.1 at the 75th percentile.

    View Source
  13. 13.

    Ellis, P. (OneRedOak). claude-code-workflows: design-review agent and CLAUDE.md snippet using Playwright MCP at 1440, 768 and 375 px.

    View Source
  14. 14.

    Parks, R.A. (10 Feb 2026). Giving Claude Code Eyes with Playwright MCP. ap7i.com.

    View Source
Share:
Get in touch

Want this for your brand?

Read something here you would like running in your business? Tell us the goal and we will send a plan and a price.