Skip to content

feat(tutorials): add AI agent performance guide - #19487

Open
willwearing wants to merge 6 commits into
masterfrom
agent/ai-agent-performance-blog
Open

feat(tutorials): add AI agent performance guide#19487
willwearing wants to merge 6 commits into
masterfrom
agent/ai-agent-performance-blog

Conversation

@willwearing

@willwearing willwearing commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

Summary

  • Adds a tutorial for measuring and improving AI agent performance with traces, clusters, sentiment, evaluations, dashboards, workflows, and experiments.
  • Publishes the guide under /tutorials/understand-ai-agent-performance.
  • Links each step to the current PostHog documentation.
  • Adds the AI product improvement loop graphic.
  • Adds all 10 product screenshots from the source document at their matching sections.
  • Adds Will to the site author data.

Source draft: https://docs.google.com/document/d/1HUxNDJ8JyfuJFzEy9ExzWzvmcoLYwjhh7dXTEoxA1yw/edit?tab=t.8sy311z29cgt

Checks

  • Parsed and validated the tutorial frontmatter.
  • Compiled the tutorial with the repo's MDX dependency.
  • Ran Markdown lint with zero errors.
  • Verified all 11 referenced image paths exist.
  • Parsed and formatted authors.json.
  • Ran git diff --check.

@github-actions github-actions Bot added the blog label Aug 14, 2026
@github-actions

github-actions Bot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

Deploy preview

Status Details Updated (UTC)
🟢 Ready View preview Aug 14, 2026 04:12PM

@github-actions

github-actions Bot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

Vale prose linter → found 0 errors, 4 warnings, 0 suggestions in your markdown

Full report → Copy the linter results into an LLM to batch-fix issues.

Linter being weird? Update the rules!

contents/tutorials/understand-ai-agent-performance.mdx — 0 errors, 4 warnings, 0 suggestions
Line Severity Message Rule
151:34 warning Capitalize 'Session Replay' for PostHog's product. Use 'session replay' for the general industry concept. PostHogBase.ProductNames
155:12 warning Capitalize 'Session Replay' for PostHog's product. Use 'session replay' for the general industry concept. PostHogBase.ProductNames
339:101 warning Capitalize 'Experiments' for PostHog's product. Use 'experiments' for the general industry concept. PostHogBase.ProductNames
376:143 warning 'suuucckk' is a possible misspelling. PostHogBase.Spelling

@github-actions

github-actions Bot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

Bundle report

Total JS (gzip)

7.55 MiB (no change)

Eager graph (modules shipped in each entrypoint's initial chunks)

Entrypoint Eager size Budget Modules
app 16.90 MiB (no change) report-only 2018
Largest modules in the app closure
Module Size
./src/data/mcp-tools.json 990.3 KiB
css ./node_modules/.pnpm/css-loader@5.2.7_webpack@5.101.3/node_modules/css-loader/dist/cjs.js??ruleSet[1].rules[8].oneOf[1].use[1]!./node_modules/.pnpm/postcss-loader@4.3.0_postcss@8.5.6_webpack@5.101.3/node_modules/postcss-loader/dist/cjs.js??ruleSet[1].rules[8].oneOf[1].use[2]!./src/styles/global.css 746.7 KiB
./src/components/Stickers/Stickers.tsx 696.4 KiB
./node_modules/.pnpm/@radix-ui+react-icons@1.3.2_react@18.3.1/node_modules/@radix-ui/react-icons/dist/react-icons.esm.js 481.4 KiB
./node_modules/.pnpm/rehype-raw@7.0.0/node_modules/rehype-raw/lib/index.js + 29 modules 395.1 KiB
./node_modules/.pnpm/@posthog+icons@0.36.6_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.cjs.js 364.8 KiB
./node_modules/.pnpm/@posthog+icons@0.36.6_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js 354.8 KiB
./src/hooks/useCustomers.tsx + 54 modules 354.5 KiB
./node_modules/.pnpm/react-markdown@8.0.7_@types+react@16.14.66_react@18.3.1/node_modules/react-markdown/lib/react-markdown.js + 88 modules 351.4 KiB
./src/components/ProductComparisonTable/index.tsx + 126 modules 301.4 KiB
./node_modules/.pnpm/cloudinary-core@2.14.0_lodash@4.17.21/node_modules/cloudinary-core/cloudinary-core.js 281.9 KiB
./src/components/SearchUI/index.tsx + 87 modules 273.0 KiB
./node_modules/.pnpm/@posthog+brand@0.8.0_react@18.3.1/node_modules/@posthog/brand/dist/generated/hoggies/svg/magnifying-glass.mjs 254.7 KiB
./node_modules/.pnpm/framer-motion@10.18.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/framer-motion/dist/es/render/dom/motion.mjs + 109 modules 253.9 KiB
./node_modules/.pnpm/d3@7.9.0/node_modules/d3/src/index.js + 208 modules 247.4 KiB

Eager-graph budgets are report-only until a baseline is established. Sizes are gzip of public/**/*.js; eager size is webpack module source bytes for the modules actually shipped in the entrypoint's initial chunks (post-tree-shake).

@willwearing willwearing changed the title feat(blog): add AI agent performance guide feat(tutorials): add AI agent performance guide Aug 14, 2026
@willwearing
willwearing marked this pull request as ready for review August 14, 2026 16:00
@willwearing
willwearing requested a review from a team as a code owner August 14, 2026 16:00

@jina-yoon jina-yoon left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hey will, this is looking good! just two fixes:

  1. for images, please follow the instructions here on how to upload assets for the website :)

  2. the > block quotes in the blog are really big so should be used sparingly. could you remove some of them and only keep like 2-3 for the most important callout quotes?

metaDescription: A practical guide to using traces, clusters, sentiment, evaluations, dashboards, and workflows to measure and improve an AI agent.
---

If you're anything like literally every other tech company in the world right now, chances are you've already shipped, or are planning to ship, an AI product. Tools of the past make it hard (impossible?!) to understand how your product is performing. Users might return every day because your agent works brilliantly, or because it tells them the weather, which isn't that useful if you're a finance company...

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
If you're anything like literally every other tech company in the world right now, chances are you've already shipped, or are planning to ship, an AI product. Tools of the past make it hard (impossible?!) to understand how your product is performing. Users might return every day because your agent works brilliantly, or because it tells them the weather, which isn't that useful if you're a finance company...
If you're anything like literally every other tech company in the world right now, chances are you've already shipped, or are planning to ship, an AI product. Tools of the past make it hard (impossible, really) to understand how your product is performing. Users might return every day because your agent works brilliantly, or because it tells them the weather, which isn't that useful if you're a finance company.


If you're anything like literally every other tech company in the world right now, chances are you've already shipped, or are planning to ship, an AI product. Tools of the past make it hard (impossible?!) to understand how your product is performing. Users might return every day because your agent works brilliantly, or because it tells them the weather, which isn't that useful if you're a finance company...

This article is my attempt at tackling the real questions: When someone asks your agent to do something, does it actually succeed? And if it doesn't, how the hell do you improve it in a measurable way?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
This article is my attempt at tackling the real questions: When someone asks your agent to do something, does it actually succeed? And if it doesn't, how the hell do you improve it in a measurable way?
This tutorial is about answering the real questions that matter when you ship an AI agent: When someone asks your agent to do something, does it actually succeed? And if it doesn't, how the hell do you improve it in a measurable way?


One person might ask:

> How much money do I have in my bank account?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

these quotes look a little too large in the preview, do they need to be in this format?

Image


One simple starting point is a [boolean (true or false) eval](/docs/ai-evals):

> Did the agent successfully answer the user's question? True or false.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

same here with the quote format!


"Was this a good response?" is less useful than:

> Did the agent return the correct bank balance using the customer's live account data? True or false.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

formatting here too


Another person might ask:

> I want to understand how much money I have in my account today.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Image


It's doing for UX research what LLMs have done for development speed.

Which is good, because otherwise we're all going to spend the next five years watching videos of people arguing with chatbots, and that would suuucckk!

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
Which is good, because otherwise we're all going to spend the next five years watching videos of people arguing with chatbots, and that would suuucckk!
Which is good, because otherwise we're all going to spend the next five years watching videos of people arguing with chatbots. And that would suuuck.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants