feat(tutorials): add AI agent performance guide - #19487
Conversation
Deploy preview
|
|
Vale prose linter → found 0 errors, 4 warnings, 0 suggestions in your markdown Full report → Copy the linter results into an LLM to batch-fix issues. Linter being weird? Update the rules!
|
| Line | Severity | Message | Rule |
|---|---|---|---|
| 151:34 | warning | Capitalize 'Session Replay' for PostHog's product. Use 'session replay' for the general industry concept. | PostHogBase.ProductNames |
| 155:12 | warning | Capitalize 'Session Replay' for PostHog's product. Use 'session replay' for the general industry concept. | PostHogBase.ProductNames |
| 339:101 | warning | Capitalize 'Experiments' for PostHog's product. Use 'experiments' for the general industry concept. | PostHogBase.ProductNames |
| 376:143 | warning | 'suuucckk' is a possible misspelling. | PostHogBase.Spelling |
Bundle reportTotal JS (gzip)7.55 MiB (no change) Eager graph (modules shipped in each entrypoint's initial chunks)
Largest modules in the
|
| Module | Size |
|---|---|
./src/data/mcp-tools.json |
990.3 KiB |
css ./node_modules/.pnpm/css-loader@5.2.7_webpack@5.101.3/node_modules/css-loader/dist/cjs.js??ruleSet[1].rules[8].oneOf[1].use[1]!./node_modules/.pnpm/postcss-loader@4.3.0_postcss@8.5.6_webpack@5.101.3/node_modules/postcss-loader/dist/cjs.js??ruleSet[1].rules[8].oneOf[1].use[2]!./src/styles/global.css |
746.7 KiB |
./src/components/Stickers/Stickers.tsx |
696.4 KiB |
./node_modules/.pnpm/@radix-ui+react-icons@1.3.2_react@18.3.1/node_modules/@radix-ui/react-icons/dist/react-icons.esm.js |
481.4 KiB |
./node_modules/.pnpm/rehype-raw@7.0.0/node_modules/rehype-raw/lib/index.js + 29 modules |
395.1 KiB |
./node_modules/.pnpm/@posthog+icons@0.36.6_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.cjs.js |
364.8 KiB |
./node_modules/.pnpm/@posthog+icons@0.36.6_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js |
354.8 KiB |
./src/hooks/useCustomers.tsx + 54 modules |
354.5 KiB |
./node_modules/.pnpm/react-markdown@8.0.7_@types+react@16.14.66_react@18.3.1/node_modules/react-markdown/lib/react-markdown.js + 88 modules |
351.4 KiB |
./src/components/ProductComparisonTable/index.tsx + 126 modules |
301.4 KiB |
./node_modules/.pnpm/cloudinary-core@2.14.0_lodash@4.17.21/node_modules/cloudinary-core/cloudinary-core.js |
281.9 KiB |
./src/components/SearchUI/index.tsx + 87 modules |
273.0 KiB |
./node_modules/.pnpm/@posthog+brand@0.8.0_react@18.3.1/node_modules/@posthog/brand/dist/generated/hoggies/svg/magnifying-glass.mjs |
254.7 KiB |
./node_modules/.pnpm/framer-motion@10.18.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/framer-motion/dist/es/render/dom/motion.mjs + 109 modules |
253.9 KiB |
./node_modules/.pnpm/d3@7.9.0/node_modules/d3/src/index.js + 208 modules |
247.4 KiB |
Eager-graph budgets are report-only until a baseline is established. Sizes are gzip of public/**/*.js; eager size is webpack module source bytes for the modules actually shipped in the entrypoint's initial chunks (post-tree-shake).
jina-yoon
left a comment
There was a problem hiding this comment.
hey will, this is looking good! just two fixes:
-
for images, please follow the instructions here on how to upload assets for the website :)
-
the > block quotes in the blog are really big so should be used sparingly. could you remove some of them and only keep like 2-3 for the most important callout quotes?
| metaDescription: A practical guide to using traces, clusters, sentiment, evaluations, dashboards, and workflows to measure and improve an AI agent. | ||
| --- | ||
|
|
||
| If you're anything like literally every other tech company in the world right now, chances are you've already shipped, or are planning to ship, an AI product. Tools of the past make it hard (impossible?!) to understand how your product is performing. Users might return every day because your agent works brilliantly, or because it tells them the weather, which isn't that useful if you're a finance company... |
There was a problem hiding this comment.
| If you're anything like literally every other tech company in the world right now, chances are you've already shipped, or are planning to ship, an AI product. Tools of the past make it hard (impossible?!) to understand how your product is performing. Users might return every day because your agent works brilliantly, or because it tells them the weather, which isn't that useful if you're a finance company... | |
| If you're anything like literally every other tech company in the world right now, chances are you've already shipped, or are planning to ship, an AI product. Tools of the past make it hard (impossible, really) to understand how your product is performing. Users might return every day because your agent works brilliantly, or because it tells them the weather, which isn't that useful if you're a finance company. |
|
|
||
| If you're anything like literally every other tech company in the world right now, chances are you've already shipped, or are planning to ship, an AI product. Tools of the past make it hard (impossible?!) to understand how your product is performing. Users might return every day because your agent works brilliantly, or because it tells them the weather, which isn't that useful if you're a finance company... | ||
|
|
||
| This article is my attempt at tackling the real questions: When someone asks your agent to do something, does it actually succeed? And if it doesn't, how the hell do you improve it in a measurable way? |
There was a problem hiding this comment.
| This article is my attempt at tackling the real questions: When someone asks your agent to do something, does it actually succeed? And if it doesn't, how the hell do you improve it in a measurable way? | |
| This tutorial is about answering the real questions that matter when you ship an AI agent: When someone asks your agent to do something, does it actually succeed? And if it doesn't, how the hell do you improve it in a measurable way? |
|
|
||
| One person might ask: | ||
|
|
||
| > How much money do I have in my bank account? |
|
|
||
| One simple starting point is a [boolean (true or false) eval](/docs/ai-evals): | ||
|
|
||
| > Did the agent successfully answer the user's question? True or false. |
There was a problem hiding this comment.
same here with the quote format!
|
|
||
| "Was this a good response?" is less useful than: | ||
|
|
||
| > Did the agent return the correct bank balance using the customer's live account data? True or false. |
|
|
||
| Another person might ask: | ||
|
|
||
| > I want to understand how much money I have in my account today. |
|
|
||
| It's doing for UX research what LLMs have done for development speed. | ||
|
|
||
| Which is good, because otherwise we're all going to spend the next five years watching videos of people arguing with chatbots, and that would suuucckk! |
There was a problem hiding this comment.
| Which is good, because otherwise we're all going to spend the next five years watching videos of people arguing with chatbots, and that would suuucckk! | |
| Which is good, because otherwise we're all going to spend the next five years watching videos of people arguing with chatbots. And that would suuuck. |


Summary
/tutorials/understand-ai-agent-performance.Source draft: https://docs.google.com/document/d/1HUxNDJ8JyfuJFzEy9ExzWzvmcoLYwjhh7dXTEoxA1yw/edit?tab=t.8sy311z29cgt
Checks
authors.json.git diff --check.