mirror of
https://github.com/microsoft/vscode.git
synced 2026-09-29 18:09:01 +01:00
* Allow chat-lib embedders to supply the tokenizer implementation Embedders of @vscode/chat-lib that already have the cl100k/o200k BPE dictionaries in memory currently load up to three copies per process (~100 MB per encoder per copy). Open both tokenizer code paths to a host-supplied implementation: - INESProviderOptions gains tokenizerProvider, following the existing optional service override pattern (languageDiagnosticsService etc.). - The completions-core tokenizer module gains setExternalTokenizerProvider(), checked by getTokenizer(), and initializeTokenizers becomes a lazy thenable so the dictionary load starts on first await (ghostText already awaits it per request) rather than at module evaluation - giving hosts a chance to register a provider before any load happens. Without a provider the only change is load timing. Fixes #321098 Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Address review: harden the external tokenizer provider hook - setExternalTokenizerProvider throws on re-registration or after the built-in load has started, freezing the load strategy. - The initializeTokenizers gate never rejects (matching the built-in best-effort contract); external loads are single-flight and a failed attempt clears the memo so a later await can retry. - getTokenizer keeps its never-throws contract: a throwing provider falls back to the module's existing fallback chain. - Use ?? instead of || for the tokenizerProvider option fallback. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Fix tokenizer tests racing the now-lazy dictionary load The cl100k/o200k mocha suites captured getTokenizer() at suite definition time, which only worked because the eager IIFE had loaded the dictionaries by then; with lazy init they captured the approximate fallback (24 failures on the Copilot test job). Fetch the tokenizer in suiteSetup after awaiting initializeTokenizers instead. getTokenizer() also kicks the lazy load on a miss now, so production callers that never await still converge on exact tokenization. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Make the completions tokenizer override a factory option Replace the process-global registration race with a factory option. `setExternalTokenizerProvider` no longer throws (first-wins, never crashes host init), and `createInlineCompletionsProvider` installs the provider synchronously before any tokenization, so there is no ordering contract for embedders to get right. Add `IInlineCompletionsProviderOptions.tokenizerProvider` and export `ExternalTokenizerProvider` / `Tokenizer` / `TokenizerName` from the chat-lib public surface so embedders can actually implement and pass the option. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Replace lazy-thenable initializeTokenizers with ensureTokenizersLoaded() The exported `initializeTokenizers` was a hand-rolled fake Promise (an object with then/catch/finally + a `Symbol.toStringTag` spoof, cast `as Promise<void>`) that re-invoked `loadTokenizers()` on every settle and lied to `instanceof Promise`. Replace it with a plain `ensureTokenizersLoaded()` function that returns the real promise from `loadTokenizers()`. Behavior is unchanged: the load still starts lazily on first call (so an embedding host can install an external provider first), is single-flighted, and never rejects. Update the one prod awaiter (ghostText) and the test call sites. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Reword getTokenizer fire-and-forget comment The comment claimed callers that never call ensureTokenizersLoaded() "still converge on exact tokenization", which overstates the guarantee: when the dictionary load fails, the approximate tokenizer fallback persists. Clarify that exact tokenization is reached on a later call only once the load succeeds. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Add unit tests for the external tokenizer provider The host-overridable tokenizer path (setExternalTokenizerProvider, the ensureTokenizersLoaded single-flight gate, and the never-throws getTokenizer fallback) had no coverage. Add a "Tokenizer Test Suite - external provider" suite with a fake provider covering delegation, concurrent single-flight, fallback-on-throw, resolve-and-retry on load failure, and first-wins registration. Because the completions tokenizer is process-global, add a test-only resetTokenizersForTest() that restores the module's initial state, called in the suite's setup and teardown so it does not leak into the other suites that share the mocha process. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Fix chat-lib extraction by re-exporting tokenizer symbols locally The chat-lib extractor (extractChatLib.ts) rewrites `import ... from` module specifiers to point at the bundled `_internal/` tree, but it does not rewrite `export ... from` re-export specifiers. The tokenizer re-exports in chatLibMain.ts used `export { TokenizerName } from '...tokenization'`, whose deep relative path was left untouched in the extracted main.ts and no longer resolved — breaking `npm run extract-chat-lib` with TS2307 and failing the chat-lib tests job on all three platforms. Follow the convention already used everywhere else in this entry point (e.g. the service re-exports just above): import the symbols at the top, where the extractor does rewrite the path, and re-export the local bindings without a `from` clause. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Mason Chen <jiec@microsoft.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>