mirror of
https://github.com/abhinavxd/libredesk.git
synced 2026-10-04 12:31:32 +00:00
b570502a20
Some OpenAI-compatible endpoints reject max_tokens and want max_completion_tokens. Before, every request paid a 400 round trip to learn this. Now we cache the swap per base URL and model, so once one request adapts, later requests send the right param up front. Only adapt when the error code is unsupported_parameter, not on any max_tokens error. Parse the usage block from provider responses into a TokenUsage struct and log prompt/completion/total tokens. Added debug logging of model content, tool results, and RAG chunk text to make agent runs easier to trace. Widget side: clean up chat bubble spacing by moving margins off the text and onto trailing elements, and drop the bottom margin on the last paragraph/list in rendered HTML so bubbles don't have extra padding.