Files
vscode/extensions/copilot/src/lib
effc60ba22 Allow chat-lib embedders to supply the tokenizer implementation (#321920)
* Allow chat-lib embedders to supply the tokenizer implementation

Embedders of @vscode/chat-lib that already have the cl100k/o200k BPE
dictionaries in memory currently load up to three copies per process
(~100 MB per encoder per copy). Open both tokenizer code paths to a
host-supplied implementation:

- INESProviderOptions gains tokenizerProvider, following the existing
  optional service override pattern (languageDiagnosticsService etc.).
- The completions-core tokenizer module gains
  setExternalTokenizerProvider(), checked by getTokenizer(), and
  initializeTokenizers becomes a lazy thenable so the dictionary load
  starts on first await (ghostText already awaits it per request)
  rather than at module evaluation - giving hosts a chance to register
  a provider before any load happens. Without a provider the only
  change is load timing.

Fixes #321098

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Address review: harden the external tokenizer provider hook

- setExternalTokenizerProvider throws on re-registration or after the
  built-in load has started, freezing the load strategy.
- The initializeTokenizers gate never rejects (matching the built-in
  best-effort contract); external loads are single-flight and a failed
  attempt clears the memo so a later await can retry.
- getTokenizer keeps its never-throws contract: a throwing provider
  falls back to the module's existing fallback chain.
- Use ?? instead of || for the tokenizerProvider option fallback.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Fix tokenizer tests racing the now-lazy dictionary load

The cl100k/o200k mocha suites captured getTokenizer() at suite
definition time, which only worked because the eager IIFE had loaded
the dictionaries by then; with lazy init they captured the approximate
fallback (24 failures on the Copilot test job). Fetch the tokenizer in
suiteSetup after awaiting initializeTokenizers instead. getTokenizer()
also kicks the lazy load on a miss now, so production callers that
never await still converge on exact tokenization.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Make the completions tokenizer override a factory option

Replace the process-global registration race with a factory option.
`setExternalTokenizerProvider` no longer throws (first-wins, never crashes
host init), and `createInlineCompletionsProvider` installs the provider
synchronously before any tokenization, so there is no ordering contract for
embedders to get right.

Add `IInlineCompletionsProviderOptions.tokenizerProvider` and export
`ExternalTokenizerProvider` / `Tokenizer` / `TokenizerName` from the chat-lib
public surface so embedders can actually implement and pass the option.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Replace lazy-thenable initializeTokenizers with ensureTokenizersLoaded()

The exported `initializeTokenizers` was a hand-rolled fake Promise (an object
with then/catch/finally + a `Symbol.toStringTag` spoof, cast `as Promise<void>`)
that re-invoked `loadTokenizers()` on every settle and lied to `instanceof
Promise`. Replace it with a plain `ensureTokenizersLoaded()` function that
returns the real promise from `loadTokenizers()`.

Behavior is unchanged: the load still starts lazily on first call (so an
embedding host can install an external provider first), is single-flighted, and
never rejects. Update the one prod awaiter (ghostText) and the test call sites.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Reword getTokenizer fire-and-forget comment

The comment claimed callers that never call ensureTokenizersLoaded() "still
converge on exact tokenization", which overstates the guarantee: when the
dictionary load fails, the approximate tokenizer fallback persists. Clarify
that exact tokenization is reached on a later call only once the load succeeds.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Add unit tests for the external tokenizer provider

The host-overridable tokenizer path (setExternalTokenizerProvider, the
ensureTokenizersLoaded single-flight gate, and the never-throws getTokenizer
fallback) had no coverage. Add a "Tokenizer Test Suite - external provider"
suite with a fake provider covering delegation, concurrent single-flight,
fallback-on-throw, resolve-and-retry on load failure, and first-wins
registration.

Because the completions tokenizer is process-global, add a test-only
resetTokenizersForTest() that restores the module's initial state, called in
the suite's setup and teardown so it does not leak into the other suites that
share the mocha process.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Fix chat-lib extraction by re-exporting tokenizer symbols locally

The chat-lib extractor (extractChatLib.ts) rewrites `import ... from` module
specifiers to point at the bundled `_internal/` tree, but it does not rewrite
`export ... from` re-export specifiers. The tokenizer re-exports in
chatLibMain.ts used `export { TokenizerName } from '...tokenization'`, whose
deep relative path was left untouched in the extracted main.ts and no longer
resolved — breaking `npm run extract-chat-lib` with TS2307 and failing the
chat-lib tests job on all three platforms.

Follow the convention already used everywhere else in this entry point (e.g.
the service re-exports just above): import the symbols at the top, where the
extractor does rewrite the path, and re-export the local bindings without a
`from` clause.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Mason Chen <jiec@microsoft.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-06-18 15:28:15 +00:00
..