Skip to content

fix(doctor): stop three readiness probes leaking non-fatal errors to console#351

Merged
placerda merged 1 commit into
developfrom
fix/doctor-console-noise
Jul 16, 2026
Merged

fix(doctor): stop three readiness probes leaking non-fatal errors to console#351
placerda merged 1 commit into
developfrom
fix/doctor-console-noise

Conversation

@placerda

Copy link
Copy Markdown
Contributor

What

agentops doctor was leaking three low-level, non-fatal probe errors straight to the console during readiness checks. Each probe already degrades gracefully; only the reporting was noisy. This stops all three from printing.

1. LLM judge -> HTTP 404

WARNING: llm_assist: judge call failed: Error code: 404

The judge used the Foundry project OpenAI client on the legacy /openai/ route, which returns 404 for chat completions. The client is now normalized to the stable /openai/v1/ route (clearing the injected api-version query param) via client.with_options(...), reusing the same credential and token refresh. Non-Foundry endpoints are returned unchanged. The judge now runs instead of 404ing.

2. OpenAI data-plane RBAC check -> UnsupportedQuery

(UnsupportedQuery) The filter 'atScopeAndAbove() and assignedTo(...)' is not supported.

The check built a combined ARM role-assignment filter that ARM rejects. It now sends the supported assignedTo('<oid>') filter, which already returns assignments at the target scope and every ancestor scope. The check runs instead of skipping with a console error.

3. App Insights probe -> read timeout

INFO: Rate-limit App Insights probe failed (non-fatal): The read operation timed out

_query_application_insights caught only urllib.error.URLError, so a read timeout (socket.timeout, an OSError but not a URLError) escaped and printed. It now catches OSError (covering both) and the request timeout was raised from 10s to 30s. Slow App Insights degrades quietly.

Tests

  • New coverage for the OpenAI client normalization (_normalize_foundry_openai_client), the corrected RBAC filter, and the App Insights timeout/URLError degrade-to-None path.
  • Full suite: 1101 passed, 1 skipped.

Notes

  • Bug fix, no linked issue required (per CONTRIBUTING).
  • Targets develop.
  • CHANGELOG updated under ## [Unreleased].

…console

- LLM judge: normalize the Foundry project OpenAI client to the /openai/v1/
  route (clearing the injected api-version) so chat completions stop returning
  HTTP 404. Reuses the same credential/token refresh; non-Foundry endpoints are
  left unchanged.
- OpenAI data-plane RBAC check: send the supported assignedTo('<oid>') ARM
  filter instead of the combined atScopeAndAbove()+assignedTo() form that ARM
  rejects with UnsupportedQuery. assignedTo alone already covers scope plus
  ancestor scopes.
- App Insights probe: catch OSError (covers socket.timeout, which is not a
  URLError) and raise the request timeout from 10s to 30s, so a slow App
  Insights degrades quietly instead of printing a non-fatal INFO line.

Adds unit coverage for the OpenAI client normalization, the corrected RBAC
filter, and the App Insights timeout/URLError degrade-to-None path.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: af408082-0971-42ea-828c-8155d4110101
@placerda
placerda merged commit fa9e48b into develop Jul 16, 2026
12 checks passed
@placerda
placerda deleted the fix/doctor-console-noise branch July 16, 2026 11:12
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant