How autopilot gathers evidence

Produce a report · 07

How autopilot gathers evidence

Understand how autopilot collects connected reports, page speed and accessibility checks for the pages a journey reached, browser observations, screenshots, and Microsoft Clarity behaviour evidence.

10 min read Updated 2 October 2026

01 / Connected reports

Let autopilot collect the connected evidence for the saved period

Connect GA4, and Search Console if you use it, at organisation level, then save the journey’s evidence period and select Start autopilot. The evidence checkpoint explains what each source adds and what the report cannot measure without it. GA4 is required: every finding in a journey is ranked by it. The checkpoint also asks whether this website has heatmaps. If you answer yes, the desktop AI adds Microsoft Clarity collection to the journey; if you answer no, the journey continues without it. There is no manual heatmap upload step. The server collects and saves the connected reports as the first resumable stage of the run.

  1. Autopilot collects the GA4 report set for the exact saved period, including coverage and data-quality metadata.
  2. When Search Console is connected, it collects web search plus image, video, Discover, Google News, and News evidence when Google returns rows, with page-and-query detail for detected search appearances.
  3. Each source is saved as a durable task. A temporary API or network failure can be retried without repeating completed collection.
  4. If the connection or property is wrong, correct it under Data Sources before starting a fresh run.

Search Console and Clarity are optional. Without them the journey continues and the report states which insights could not be calculated. Google can return only the top rows for large Search Console queries, so a report that reaches its row limit is marked as potentially incomplete. See Journey analyses for how each source is used in the report.

02 / Page speed

Measure real-user performance on the pages the journey reached

When the journey fieldwork finishes, UX Robot’s server measures the pages the journey actually reached, on the locked device, while the desktop agent carries on. Core Web Vitals comes from the Chrome UX Report API: current URL and origin field data plus origin history, with the exact collection periods. Third-party pages, such as a payment provider at the end of a journey, are not measured. An unavailable CrUX snapshot means Google does not have enough eligible real-user data for that target and device. It is an evidence gap, not a passing result.

Understand the p75 values

p75 means the 75th percentile. Seventy-five per cent of qualifying visits performed at or better than the reported value, while the remaining 25 per cent were worse. Lower values are better for all three Core Web Vitals.

Swipe or scroll the table horizontally to compare every threshold.

Core Web Vitals meanings and p75 thresholds
Metric What it measures Good Needs improvement Poor
LCP p75 Largest Contentful Paint: loading performance, measured by how long the page’s main visible content takes to appear. 2.5 seconds or less Over 2.5 to 4 seconds Over 4 seconds
INP p75 Interaction to Next Paint: responsiveness, measured by how quickly the page visibly responds after a click, tap, or keyboard interaction. 200 ms or less Over 200 to 500 ms Over 500 ms
CLS p75 Cumulative Layout Shift: visual stability, measured by how much visible content unexpectedly moves while the page is being used. CLS is a score, not a duration. 0.10 or less Over 0.10 to 0.25 Over 0.25

A target passes Core Web Vitals only when its LCP, INP, and CLS are all in the good range at p75 for the relevant device group. For example, an LCP p75 of 3.4 seconds means 75 per cent of qualifying visits loaded the main visible content within 3.4 seconds, 25 per cent took longer, and the result needs improvement.

03 / Accessibility and performance

Check the journey’s pages for barriers, bad practice and slow loading

Accessibility problems are usability problems. A form field with no label, text that does not contrast with its background, or a link that reads only “click here” will stop some people completing a journey, and analytics can show you the drop-off without ever telling you the cause.

UX Robot runs automated Lighthouse accessibility, best-practice, and performance checks through PageSpeed Insights on the same reached pages. Results are grouped by problem rather than by page, with the pages and example elements affected. Google renders each page before scoring it, so the checks run in the background; a page that times out is tried again automatically, and one that still cannot be checked is marked as failed while the rest complete. On a signed-in journey only field data is collected, because Lighthouse cannot use your session. A page tested before it ships is checked by the desktop agent on the engineer’s computer instead, because Google cannot reach a test address.

The same checks return performance diagnostics: which stylesheet or script blocked rendering and for how long, which images are larger than the space they fill, and how the largest content paint divides between server response, resource discovery, download and rendering. These are laboratory measurements from one rendered page load, not what your visitors experienced. The report has a dedicated Performance section that separates Core Web Vitals field status from Lighthouse laboratory coverage, including not collected, unavailable, failed and still-running checks. Lighthouse performance diagnostics are supporting laboratory evidence and are labelled that way.

The desktop agent also checks practical accessibility during the journey: keyboard access, focus, labels, error messages and recovery where it is safely available. Each issue lists the pages it appears on and example elements such as input#email, so the work can be handed to a designer or engineer without them repeating the check.

04 / Behaviour

Let the desktop agent gather Microsoft Clarity evidence

At the evidence checkpoint, answer Yes — collect them from Microsoft Clarity when heatmaps are available. After the browser fieldwork establishes the journey’s URLs, the desktop agent reviews relevant click, scroll, attention, and recording evidence for those pages. It saves the page, device, date filters, sample context, observed pattern, and limitations instead of asking the engineer to upload heatmaps manually.

If Microsoft Clarity is signed out when that collection step begins, Autopilot saves its place and asks the engineer to sign in. If you answer that no heatmaps are available, UX Robot does not add a Clarity task or show Clarity instructions; the journey continues and the report explains the behavioural evidence limitation.

Browser inspection shows what the interface does; Clarity shows how observed visitors used it. The report keeps those evidence types separate and does not turn a small recording sample into a claim about the whole audience.

05 / Observations

Save browser observations with clear provenance

Each desktop task saves structured observations in bounded chunks. An observation records the URL, page area, device or viewport, action, expected and observed result, applicable UX principle, confidence, limitation, and a location screenshot. Every failed visual or interface observation must include a screenshot of the exact tested viewport and the affected element bounds so UX Robot can add the highlight deterministically. Highlighted finding images in the report remain the captured screenshot with that overlay rather than an AI-generated recreation.

The agent must not submit personal, authentication, payment, or confidential information. When a useful screenshot would contain sensitive content, show a signed-in third-party tool, describe an issue that cannot be shown in one viewport, or cannot be captured with the available browser, it must save the explicit reason, record the location in words, and leave the image out. The server rejects a failed observation that silently omits both the capture and the reason.

When the screenshot already exists as a local JPEG, PNG, or WebP file, the agent marks the capture as pending and checkpoints the observation. It then uploads the binary file through a short-lived one-use link by running the supplied command in its shell. This avoids copying large Base64 image data through the MCP conversation. UX Robot can verify the local SHA-256, confirms the stored checksum, and refuses to finish the task until every pending upload succeeds.

One observation records one interface state and one screenshot. A second state is saved as supplementary evidence rather than replacing the first. While fieldwork is active, the desktop agent can also correct factual observation text; UX Robot retains the correction reason in the evidence history.

Still stuck?

Use the contact form and include the page, action, and exact error message involved. Contact UX Robot.