Real LLM chat replacing the mock, backend-centric per the plan:
- Multi-provider API-key mode (OpenAI / Anthropic / Gemini via the AI SDK) plus
local Ollama (OpenAI-compatible endpoint). Provider is derived from the
picked model id; the matching stored key is used. New user_ai_settings table
holds per-user config with provider API keys encrypted at rest (AES-256-GCM,
src/lib/crypto.ts, keyed by AI_CREDENTIALS_KEY).
- POST /api/ai/config (get/put, secrets never returned), POST /api/ai/test
(Ollama ping / key presence), POST /api/ai/import (approved migration commit,
re-validated server-side, reuses the audited patient service).
- POST /api/chat: streamText agent with tools (getPatient, getPatientLabs,
searchPatients, previewImport). Real record data streams to the clinician as
custom data parts (cards) while the model sees only Veil-redacted results.
- Veil (src/services/ai/veil.ts): de-identifies patient identifiers to tokens
before external calls, resolves tokens on tool args, and rehydrates the final
answer. Bypassed for local Ollama. External mode runs non-streamed so the
rehydrated text is correct. Every call is audited (provider + Veil level).
- Shared role-scoping helpers extracted to src/lib/role-scope.ts (reused by the
patient routes and chat tools so visibility rules match).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>