Commit Graph
61 Commits
Author SHA1 Message Date
Ross A. WollmanandCopilot 9b2d89997d agentHost: honor identity capture policy in host telemetry
Apply managed identity capture to host metadata and trace storage/export. Preserve existing content-capture and export-destination precedence while resolving the new identity preference against shell environment values.

Refs #336814.

🤖 Authored with GitHub Copilot on behalf of @rwoll.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-09-30 21:36:04 -07:00
Joaquín RualesandCopilot App f20366e43e Avoid stale OTel reload prompts after policy recovery (#338732)
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
2026-09-29 12:31:33 -07:00
Joaquín RualesandCopilot App e4f7a6f2f8 Handle settled blocked policy in Local OTel recovery (#338645)
* Fix Local OTel recovery after blocked policy refresh

Refs #338630. Wait for authoritative account-policy settlement instead of treating restrictive exporter values as transient.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

* Await gate-owned account policy initialization

Wait for the post-account managed-settings update before settling Local OTel policy recovery. Cover real initialization ordering, unchanged terminal gates, overlapping updates, and initialization failures.

Refs #338645.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
2026-09-29 10:52:16 +00:00
Joaquín RualesandCopilot App d850a1570c Fix managed OTel restart recovery (#338623)
* Fix managed OTel restart recovery

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

* Address managed OTel review feedback

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

* Bound managed OTel recovery and defer transient policy restarts

Refine recovery guards and policy readiness in #338623.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

* Handle forced-refresh defaults in Local OTel recovery

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

* Ignore transient Local OTel policy placeholders

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
2026-09-29 02:34:44 -07:00
Vijay UpadyaandCopilot bdc5ebe5a7 Export chat user-perceived time to first progress through OTel (#337833)
* Export chat user-perceived time to first progress through OTel

* Feedback update

* Fix formatting of chat timing protocol test

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-09-25 00:32:00 +00:00
Ross Wollman f2c57dbe11 Add governed OTel identity capture for Local agents (#336814)
Resolves #302930. Target: **1.140.0**. Builds on #336701.

## Summary

Adds opt-in, enterprise-governed identity capture to **Local Copilot Chat**:
- `user.name` on agent invocation spans, including subagents and inline agents.
- `process.user.name` and `host.name` on resources.

Identity is **off by default and independent of content capture**. Explicit managed policy wins over environment and personal preferences. Denial strips identity from subsequent exports, including queued data and SQLite, and stays latched until reload. Explicit resource attributes cannot bypass it.

Managed telemetry selects one whole block in priority order: native MDM > server > file. Omitted fields are not inherited from lower-priority sources. Resource batching is preserved; unavailable OS usernames do not stop telemetry.

**Scope:** Local only. Agent Host endpoint routing and native identity validation are tracked in #337413.

## Validation

- 314 focused tests, extension type-check, and targeted lint passed.
- Live Local checks passed: identity/content enabled, content disabled, and identity denied without reload. Re-enable/reload behavior has automated coverage, not live validation.
- Opus 5.5 re-review found no remaining issues; renewed signed approval is pushed.
- Latest CI status is tracked in the checks below.
2026-09-23 10:26:06 +02:00
Joaquín RualesandCopilot 82f629a453 Fix Copilot OTel file span serialization (#337373)
Serialize public SDK span data rather than circular processor state. Report serialization failures without writing placeholders, preserving existing successful log and metric output.

Fixes #319993

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-09-22 17:57:02 -07:00
Ross A. WollmanandCopilot 9ff9a50764 Log successful enterprise OTel recovery without another toast
Remove the redundant success notification while retaining the persisted restart acknowledgement, output log, monitoring indicator, and manual reload warnings. Add regression coverage for silent success and update the monitoring guide.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-09-18 00:12:54 -07:00
Ross A. WollmanandCopilot 5dce72ab2d Clear enterprise OTel restart notice with its extension host
Use a lifecycle-bound progress notification instead of a persistent warning so successful restart cannot leave contradictory notices. Preserve the existing restart guard and fallback grace, and cover command failure, fallback, and unmanaged settings in contribution tests.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-09-17 23:56:12 -07:00
Ross A. WollmanandCopilot 5aeae1f614 Restrict OTel recovery to newly arriving enterprise policy
Track recognized managed settings independently of export enablement so later updates to a disabled startup policy only offer reload. Add regression coverage and align OTel Settings descriptions with the retained environment-variable precedence.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-09-17 23:03:34 -07:00
Ross A. WollmanandCopilot 4771098e0d Checkpoint restart-based enterprise OTel recovery
Preserve the event-driven, single-attempt recovery implementation and OTel settings-block replacement before evaluating an in-process alternative. Refs #336102.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-09-17 22:00:08 -07:00
Zhichao Li 751c53b022 docs: align OTel guidance with Agent Host architecture 2026-09-08 17:21:12 -07:00
Osvaldo OrtegaandCopilot cd5814553e Remove all backend v1 code for Copilot cloud sessions provider (#330007)
* Agent Host changes for osortega/agents/remove-all-the-code-for-backend-v1-of

* Preserve pending cloud tasks until workspace opens

Only consume cross-window task deep links after Git reports the matching repository, preventing the source window from clearing shared pending state.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Migrate cloud session state from pull request URIs

PR-backed cloud tasks are listed under a stable /task/<id> resource and report the /<prNumber> URI they were previously listed under, so archived, pinned and read state migrates forward instead of being orphaned.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-08-10 18:36:38 -07:00
roblourens 78a7b6c291 agentHost: remove extension-host Claude implementation (#329697)
* Remove extension-host Claude implementation

Make Agent Host the sole Claude sessions implementation and remove the obsolete provider preference settings, extension runtime, SDK dependency, commands, tests, and integration glue.\n\n(Written by Copilot)\n\nCo-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Remove extension-host Claude smoke tests

Delete smoke cases and commands that target the removed extension-host Claude implementation.\n\n(Written by Copilot)\n\nCo-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Fix Agent Host Claude permissions link

Keep the Claude-specific permission documentation link for Agent Host Claude after removing the extension-host implementation.\n\n(Written by Copilot)\n\nCo-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-08-08 03:55:37 +00:00
Zhichao LiandCopilot a4a3e6c1bd otel: address PR review feedback on session-id correlation
- Rewrite BYOK session-id tests to exercise the real Anthropic/Gemini
  providers with a capturing OTel service and a correlated CapturingToken,
  instead of reproducing the attribute spread (which could not catch a
  production regression). Removes the synthetic tests from
  byokProviderSpans.spec.ts.
- agent_monitoring.md: add copilot_chat.session_id rows to the foreground
  execute_tool, Claude execute_tool and execute_hook span tables (the code
  emits it); mark chat-span gen_ai.conversation.id as conditional rather
  than required.
- agent_monitoring_arch.md: narrow the Session Correlation claim — the
  bridge only enriches the debug-panel copy, not the SDK's own OTLP export;
  document that CLI SDK-native OTLP correlation depends on the runtime
  (execute_tool covered once copilot-agent-runtime#12383 lands).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-07-09 16:37:18 -07:00
Zhichao LiandCopilot a113dd126e otel: inject gen_ai.conversation.id on CLI SDK-native spans via bridge
The Copilot CLI bridge injected copilot_chat.chat_session_id on forwarded
SDK-native spans but not the standard gen_ai.conversation.id, so
vendor-agnostic OTLP backends could not group CLI child spans (e.g.
execute_tool) by conversation without a traceId join. Inject the session
id under gen_ai.conversation.id too (without overwriting an SDK-provided
value). Also fix a type error where the execute_hook span assigned a
possibly-undefined chatSessionId to conversation.id / session_id
unconditionally.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-07-09 16:24:58 -07:00
Zhichao LiandCopilot a9dd51136d otel: document + test session-id span correlation
Add session-correlation attribute rows (gen_ai.conversation.id,
copilot_chat.session_id, copilot_chat.chat_session_id) to the chat,
execute_tool and execute_hook span tables in the monitoring docs, and a
Session Correlation section to the architecture doc. Add regression tests
locking in BYOK session-id emission and that gen_ai.conversation.id holds
the session id (not the per-turn request id).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-07-09 15:55:26 -07:00
Zhichao Li 7b0180b3c3 docs: update OTel user + agent-host docs for managed-settings precedence and new settings 2026-06-26 21:38:27 -07:00
Zhichao Li e9a25536d6 refactor: address PR review — policy precedence docs, grpc transport inference, prototype-pollution hardening 2026-06-26 21:38:26 -07:00
Zhichao Li 65bcc23a6d docs: remove internal OTel managed-settings planning notes 2026-06-26 16:58:47 -07:00
Zhichao Li 6410d9e8f5 docs: record serviceName/resourceAttributes/headers delivery in sprint 2026-06-26 16:04:32 -07:00
Zhichao Li 0fdb82fd27 docs: revise OTel managed-settings sprint plan for headers/resourceAttributes/serviceName
Records the runtime spike: the headless agent host resolves OTel from env only and
doesn't self-fetch managed telemetry, but build_resource reads OTEL_SERVICE_NAME /
OTEL_RESOURCE_ATTRIBUTES env. Revised plan delivers serviceName + resourceAttributes
to both surfaces (env for the host, programmatic for the extension) and headers to the
extension only; agent-host headers stay deferred (env would leak the token to tool
subprocesses).
2026-06-26 15:27:46 -07:00
Zhichao Li 4be6140abd docs: update OTel managed-settings plan/sprint for protocol parity 2026-06-26 13:39:17 -07:00
Zhichao Li d0ba9fcdc4 docs: add OTel managed-settings sprint plan with completion notes 2026-06-26 11:43:03 -07:00
Zhichao Li b2ca1ff9d5 docs: note policyReference type-match constraint and structured-encode location 2026-06-26 11:23:01 -07:00
Zhichao Li 1ae83f24a5 docs: add enterprise OTel managed-settings policy plan
High-level plan for VS Code enterprise control of Copilot agent-host OTel export via the cross-client telemetry managed-settings schema (matches CLI ManagedTelemetrySettings, copilot-agent-runtime #10735). Covers schema, ownership, precedence, security, delivery channels, suppressions, and touch points.
2026-06-26 11:07:36 -07:00
Eleanor Boyd 9cef07e98c Merge pull request #322002 from eleanorjboyd/quiet-otter
Add inline chat OTel invocation span
2026-06-22 11:31:17 -07:00
eleanorjboyd 349c609f15 Add inline chat OTel invocation span 2026-06-18 14:50:36 -07:00
Osvaldo OrtegaandCopilot d8716e0178 Instrument Cloud Agent backend v1/v2 rollout behind an experiment (#321818)
* Instrument Cloud Agent backend v1/v2 rollout

Put the `chat.cloudAgentBackend.version` toggle (v1 = Jobs API, v2 = Task
API) behind an experiment so the rollout can be ramped and rolled back
remotely via ExP, and add an arm-tagged telemetry surface so the two
backends emit identical funnel and guardrail signals for apples-to-apples
monitoring.

Previously v1 was partly instrumented and v2 emitted nothing, making any
A/B comparison impossible. A shared `ICloudBackendInstrumentation` is now
injected into both backends and stamps every event with `backendVersion`,
covering session creation, activation, follow-up, and per-operation
errors, plus matching OTel metrics. v2's silent catch sites now report
through it, and `TaskApiError` carries the HTTP status for error
classification. Legacy v1 backend events are consolidated into the shared
surface and kept for dashboard continuity.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Rename "arm" to "backend version" in cloud telemetry comments

"arm" is A/B-experiment jargon not used elsewhere in the codebase. Swap it
for "backend version" in comments, docstrings, and GDPR comment text to
match the existing `backendVersion` / v1 / v2 vocabulary. Comment-only
change; no identifiers or behavior affected.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Address cloud backend rollout PR review comments

- Move backend_version attribute to the canonical github.copilot.* namespace
- Derive numeric HTTP status inside operationFailed so the status measurement is populated
- Carry HTTP status on v1 (Jobs API) invalid-job errors via JobsApiError so v1 create failures classify by status
- Only emit cloud pr_ready.count for v1 (Task API activation is a first turn, not a PR)
- Add unit tests for the new cloud metrics
- Document the new cloud operation/error metrics and backend_version attribute

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Fix TaskApiBackend construction in fetchSessionList specs

Three fetchSessionList test cases constructed TaskApiBackend without the
required instrumentation argument, failing the CI typecheck (TS2554).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-06-17 16:41:16 -07:00
Zhichao Li 136b39f373 Emit gen_ai streaming OTel signals for chat responses
Adds GenAI semantic-convention streaming signals alongside the legacy
copilot_chat.time_to_first_token, tagged with gen_ai.response.model for
per-model slicing (microsoft/vscode#320651):

- gen_ai.request.stream (span attr, bool)
- gen_ai.response.time_to_first_chunk (span attr, seconds)
- gen_ai.client.operation.time_to_first_chunk (histogram metric)
- gen_ai.client.operation.time_per_output_chunk (histogram metric)

The main GitHub streaming path (chatMLFetcher) carries the full set
including per-output-chunk latency, computed from inter-chunk gaps in
FetchStreamRecorder. BYOK providers (Anthropic, Gemini) emit the stream
attr + time_to_first_chunk only, as they lack per-chunk arrival timing.
The Claude agent path inherits all signals via the shared chatMLFetcher
proxy; the Copilot CLI path emits them natively from the runtime.

Updates agent_monitoring.md and adds unit tests.
2026-06-09 17:21:11 -07:00
zhichli a09e44ba92 otel: normalize gen_ai.response.model to match request format
Anthropic responses echo the same logical model with a different separator
(e.g. request 'claude-opus-4.6' resolves to 'claude-opus-4-6'), causing
GROUP BY / DISTINCT on gen_ai.response.model to produce duplicate rows.

Add normalizeResponseModel(requestModel, responseModel) that echoes the
request value when the two strings only differ in '.' vs '-', and returns
the resolved value unchanged when it adds real specificity (e.g.
'gpt-5.4-mini' -> 'gpt-5.4-mini-2026-03-17'). Wire into the three call
sites that set gen_ai.response.model: chatMLFetcher (chat spans),
toolCallingLoop (foreground invoke_agent), and claudeOTelTracker
(Claude invoke_agent).

Fixes #318805
2026-05-28 16:06:11 -07:00
Zhichao Li b169da78ec Chat OTel: enrich Claude agent spans with github.copilot.* attributes 2026-05-26 17:26:28 -07:00
Zhichao Li ec79deaf7c Chat OTel: drop sprint plan, tighten comments and docs 2026-05-23 23:40:57 -07:00
Zhichao Li 7fc2ce6d2e docs(otel): document github.copilot.* attributes and mark copilot_chat.repo.* legacy 2026-05-22 11:15:10 -07:00
Zhichao Li 24c6799a8f docs(otel): sprint plan for github.copilot.* parity 2026-05-22 11:06:54 -07:00
Zhichao Li 6fef3dffff feat(otel): make attribute truncation configurable
Adds `github.copilot.chat.otel.maxAttributeSizeChars` setting and
`COPILOT_OTEL_MAX_ATTRIBUTE_SIZE_CHARS` env var to control truncation
of free-form OTel content attributes (prompts, responses, tool
arguments/results, hook input/output).

Default is `0` (unlimited), matching the OTel spec's
`AttributeValueLengthLimit` default of `Infinity`. Users on backends
with per-attribute size limits can set a positive value to keep OTLP
batches under the backend cap.

Plumbs the resolved limit through every call site that previously hit a
hardcoded 64KB fallback. Drops the `DEFAULT_MAX_OTEL_ATTRIBUTE_LENGTH`
constant; `truncateForOTel`'s default arg is now `0` (unlimited).

Refs #299952
2026-05-01 23:30:23 -07:00
Zhichao Li 4cd66e05d3 OTel docs: align agent_monitoring docs with current implementation
agent_monitoring.md (user-facing):
- Add github.copilot.chat.otel.dbSpanExporter.enabled to the settings table
- Add dbSpanExporter as a fourth Activation trigger
- Add Commands subsection for 'Chat: Export Agent Traces DB'
- Add claude-code row to the service.name filter table

agent_monitoring_arch.md (developer-facing):
- Correct the Multi-Agent Architecture row for Claude (synthesized from SDK
  messages via claudeMessageDispatch.ts + chatMLFetcher proxy, not 'message loop')
- Rewrite the Claude span hierarchy diagram to show subagent nesting under
  execute_tool Agent, and remove stale claudeHookRegistry/claudeCodeAgent
  references and PR-number annotations
- Replace bogus claudeCodeAgent.ts/claudeHookRegistry.ts rows in the
  Instrumentation Points table with the real emit sites
  (claudeOTelTracker.ts, claudeMessageDispatch.ts, claudeLanguageModelServer.ts,
  chatHookService.ts, geminiNativeProvider.ts)
- Expand File Structure with workspaceOTelMetadata, sessionUtils, sqlite/,
  the Claude folder layout, chatHookService, and otlpFormatConversion
- Correct the Implementations table: InMemoryOTelService is the fallback
  when OTel is disabled (not always-on alongside NodeOTelService); document
  the selection in services.ts
- Add Activation Channels subsection mapping OTelConfig.enabledVia values
- Expand the Env Var Translation table with OTEL_METRICS_EXPORTER,
  OTEL_LOGS_EXPORTER, OTEL_LOG_TOOL_DETAILS, OTEL_EXPORTER_OTLP_PROTOCOL,
  COPILOT_OTEL_EXPORTER_TYPE
- Document FilteredSpanExporter alongside DiagnosticSpanExporter in the
  Debug Panel vs OTLP Isolation section
- Update Testing tree with all current spec files
2026-04-29 16:53:15 -07:00
Zhichao Li 6e32d80c1d feat(monitoring): update cache token usage metrics for accuracy and clarity
feat(tracing): rename hook_name attribute to hook_command for better context
fix(agent): clear subagent trace contexts after request completion
2026-04-27 14:14:40 -07:00
Zhichao Li eda411ebd8 feat(monitoring): add cache token usage metrics to agent monitoring documentation 2026-04-27 14:14:37 -07:00
Zhichao Li c9807ed451 feat(telemetry): simplify Claude OTel environment configuration by removing beta tracing and content capture variables 2026-04-27 14:14:37 -07:00
Zhichao Li 4df809e53a feat(monitoring): enhance agent monitoring documentation for Claude agent tracing and usage metrics 2026-04-27 14:14:36 -07:00
Ulugbek Abdullaev 24c9ba10d2 update docs for edit capturing (#311694) 2026-04-21 09:12:02 -07:00
James Newton-King 0f31293663 Add Aspire Dashboard trace screenshot to monitoring docs 2026-04-08 02:41:14 +00:00
James Newton-King 6c656f02fa Change OTLP endpoint from gRPC to HTTP
Updated the OTLP endpoint from gRPC to HTTP and modified the related configuration settings.
2026-04-08 01:47:42 +00:00
Bhavya U 945c0c61d9 Move docs/prompts.md to .github/instructions/model-prompts.instructions.md (#4942)
- Converted standalone doc into a scoped instructions file (applyTo: src/extension/prompts/node/agent/**)
- Fixed outdated method names (resolvePrompt -> resolveSystemPrompt, PromptConstructor -> SystemPrompt)
- Added resolver interface table, DI examples, resolution order docs
- Restored concrete examples for common model misbehaviors
- Auto-loads when editing agent prompt files
2026-04-02 18:38:27 +00:00
Zhichao Li 601b3c97f6 fix: address review feedback on OTel agent activity metrics (#4801)
* fix: address review feedback on OTel agent activity metrics

* fix: guard recordEditAcceptance for accept/reject only, fix doc wording
2026-03-29 00:28:06 +00:00
Zhichao Li 05da8fb689 feat: add OTel events and metrics for agentic edit quality signals (#4794)
* docs: add OTel backfill plan for agentic change metrics

* docs: add Claude Code OTel parity analysis with feasibility + line estimates

* docs: expand plan to cover all agentic surfaces (inline chat, CLI, cloud, NES)

* docs: remove NES and Claude Code comparison, keep plan lean

* docs: add 3-pillar signal type mapping (metrics/events/traces)

* docs: re-audit signals — counters when easy, events only with useful attrs

* feat: add OTel event emitters for agentic edit quality metrics

* feat: add OTel counters and histograms for agentic edit quality metrics

* feat: wire OTel events/metrics into userActions.ts for all agentic user actions

* feat: wire OTel survival events into apply_patch, replace_string, and code_mapper tools

* feat: wire OTel counters for agent summarization and edit response metrics

* fix: resolve TypeScript errors — thread IOTelService through intent class hierarchy

* docs: update sprint plan with completion notes

* style: fix import ordering from editor auto-sort

* feat: wire OTel counters for cloud session invoke, PR ready, and CLI PR creation

* docs: consolidate OTel edit quality metrics into agent_monitoring.md

* docs: rename Edit Quality to Agent Activity & Outcome

* docs: align Edit Quality references to Agent Activity naming

* refactor: adopt Harald's type-safe metrics API (EditSource/EditOutcome, 2 survival histograms)
2026-03-28 17:13:22 +00:00
Zhichao Li d2c8aa5d67 fix: always enable content capture for CLI debug panel (#4581)
Set OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=true before
SDK init so the debug panel always shows full prompts, responses,
and tool arguments. Without this, users without the env var set
see empty content in the debug panel.

When user OTel is disabled, SDK spans go to /dev/null so captured
content never leaves the process.
2026-03-21 05:46:45 +00:00
Zhichao Li 4a4411e88e Native OTel instrumentation for Copilot CLI (background, terminal, debug panel) (#4507)
* add OTel instrumentation spec and plan for all agents

* feat: OTel instrumentation for Copilot CLI background agent

- Add agentOTelEnv.ts config derivation helpers (CLI + Claude)
- Enable SDK OtelLifecycle via env vars before LocalSessionManager ctor
- Add invoke_agent copilotcli wrapper span with traceparent propagation
- Forward OTel env vars to terminal CLI sessions
- Update spec and plan docs for all agents
- 33 tests passing (14 new + 19 existing)

* feat: filter debug-panel-only spans from OTLP export

Spans with non-standard gen_ai.operation.name values (content_event,
user_message) are excluded from external OTLP export while remaining
visible in the Agent Debug Log panel via onDidCompleteSpan.

Only GenAI-conventional operations (invoke_agent, chat, execute_tool,
embeddings, execute_hook) are exported to the user's collector.

* fix: add IOTelService to CopilotCLISessionService ctor in participant test

* fix: pass chatSessionId to CapturingToken for debug panel routing

The CapturingToken was created without chatSessionId, so the debug panel
couldn't route copilotcli OTel spans to the correct session view.

Also: Copilot CLI runtime only supports otlp-http (not gRPC). Terminal
CLI sessions require an HTTP-compatible OTLP endpoint.

* docs: add CLI HTTP-only limitation to spec and dual-port Aspire setup to test plan

* fix: forward OTel env vars to CLI terminal sessions

- Include OTel env vars in terminal profile provider path (dropdown)
  which previously only set shell info without auth/OTel env
- Pass empty env to deriveCopilotCliOTelEnv for terminal sessions so
  vars are always included regardless of process.env pollution from
  the in-process background agent
- Update test plan to use Grafana LGTM stack

* fix: add CHAT_SESSION_ID to attributes in CopilotCLISession

* docs: update OTel instrumentation specification for Copilot CLI and Claude Code

* feat: bridge SDK native OTel spans to Agent Debug panel

Replace synthetic span approach (PR #4494) with a bridge SpanProcessor
that forwards SDK-native spans from the Copilot CLI runtime's
BasicTracerProvider into the extension's IOTelService event stream.

This gives the debug panel the full SDK span hierarchy (subagents,
permissions, hooks, nested tool calls) — identical to what Grafana shows.

Architecture:
- Add injectCompletedSpan() to IOTelService interface for external span
  injection without OTLP re-export
- Create CopilotCliBridgeSpanProcessor that converts ReadableSpan to
  ICompletedSpanData, injects copilot_chat.chat_session_id from a
  traceId→sessionId map, and fires onDidCompleteSpan
- Install bridge on SDK's TracerProvider via internal
  MultiSpanProcessor._spanProcessors array (OTel SDK v2 removed the
  public addSpanProcessor API, but this internal array is the same
  pattern the SDK itself uses in forceFlush)
- Propagate traceparent from extension root span to SDK via
  otelLifecycle.updateParentTraceContext() so all spans share a traceId
- Filter bridge to only forward spans from registered CLI sessions

Code changes:
- copilotCliBridgeSpanProcessor.ts: new bridge processor
- copilotcliSession.ts: remove all synthetic spans (chat, tool, error),
  keep root invoke_agent span + traceparent propagation + bridge wiring
- copilotcliSessionService.ts: install bridge after first session
  creation, wire bridge + SDK trace context updater to sessions
- IOTelService: add injectCompletedSpan to interface + all impls
- Remove outdated synthetic span tests
- Add OTel data flow architecture diagram (HTML)

* fix: update span processing to use parent span context and enhance subagent event identification

* display names for tool call and subagent events

* docs: merge arch and spec into single developer guide

Combine agent_monitoring_arch.md (foreground-only) and agent-otel-spec.md
(all agents) into a single comprehensive developer reference covering all
four agent paths, bridge architecture, and SDK internal access warnings.

* docs: fix stale addSpanProcessor reference in data flow diagram

* chore: move plan and test docs to offline archive

These documents are reference material for the OTel sprint, not needed
in the shipped PR. Archived to ~/Documents/copilot-otel-archive/.

* test: add bridge SpanProcessor unit tests

13 tests covering: traceId filtering, parentSpanContext conversion,
CHAT_SESSION_ID injection, attribute flattening, event conversion,
HrTime→ms conversion, unregister/shutdown behavior.

* test: add span event identification and naming tests

7 tests covering invoke_agent identification logic: top-level skip,
SDK wrapper skip (no agent name), subagent detection (name attribute
and span name parsing), unknown/missing operation name handling.

* fix: always enable SDK OTel for debug panel regardless of user config

The CLI SDK's OtelLifecycle must always initialize so the bridge
processor can forward native spans to the debug panel. When user
OTel is disabled, COPILOT_OTEL_ENABLED is still set but no OTLP
endpoint is configured — the SDK creates spans (for debug panel)
but doesn't export to any external collector.

The bridge installation is also now unconditional — it installs
even when user OTel is disabled.

* chore: remove transient sprint plan

* fix: suppress SDK OTLP export when user OTel is disabled

When user OTel is disabled, force the SDK to use file exporter to
/dev/null instead of letting it default to OTLP. Also clear any
leftover OTEL_EXPORTER_OTLP_ENDPOINT from previous sessions to
prevent orphaned traces in Grafana.

* docs: add background agents section to user monitoring guide

Cover Copilot CLI (background + terminal) and Claude Code agent
tracing in the user-facing guide. Includes span hierarchy examples,
service.name filtering table, and CLI HTTP-only limitation note.

* docs: remove Claude Code from user guide (not yet supported)

* fixup! feat: OTel instrumentation for Copilot CLI background agent

* fix: address PR review comments

- Use GenAiOperationName constants in EXPORTABLE_OPERATION_NAMES (avoids drift)
- Remove unnecessary delete of OTEL_EXPORTER_OTLP_ENDPOINT from process.env
- Replace 'as any' OTel mocks with typed NoopOTelService in terminal tests
- Clarify comment on empty env arg for terminal OTel env derivation
- Add ExportResultCode.SUCCESS comment for clarity

* fixup! fix: always enable SDK OTel for debug panel regardless of user config

* fix: handle SDK native hook spans in debug panel

The SDK's OtelSessionTracker creates 'hook {type}' spans with
github.copilot.hook.type attributes (not gen_ai.operation.name).
These were silently dropped by completedSpanToDebugEvent. Now
detected by span name prefix and converted to Hook: {type} events.

* add execute_hook spans for Claude hook executions in monitoring documentation

* docs: add hook spans to CLI trace hierarchy in user guide
2026-03-20 22:53:20 +00:00
Zhichao Li 568ea9428e docs: improve OTel monitoring doc with Quick Start guide and VS Code settings examples (#4243)
- Replace env var Quick Start with step-by-step Aspire Dashboard guide
- Add concise intro explaining what the Aspire Dashboard is
- Convert all example configurations from env vars to VS Code settings JSON
- Keep env var reference table for official documentation
- Note where env vars are still required (e.g. auth headers)
2026-03-06 18:15:53 +00:00