M8B talks to its model through the OpenAI Responses API (POST /v1/responses, streaming,
function tools). AI_PROVIDER selects one of four backends. The application never branches on the
backend name; it reads the provider’s capability flags, so the columns below are exactly what the
code does. Variables are documented in CONFIGURATION.md.
AI_PROVIDER |
Backend | Conversation data goes to |
|---|---|---|
openai (default) |
Hosted OpenAI, gpt-5.6-sol |
OpenAI |
ollama |
Ollama’s /v1/responses (text model, e.g. qwen3.8:27b) |
Your Ollama host only (search queries to ollama.com if WEB_SEARCH_PROVIDER=ollama-cloud) |
vllm |
vLLM’s /v1/responses (multimodal model, e.g. Qwen 3.8 27B INT8) |
Your vLLM host only |
openai-compatible |
Any other /v1/responses server: gateway, LiteLLM, NVIDIA NIM, proxy, … |
That endpoint, wherever it is |
In every self-hosted mode the default-enabled page reader also contacts the public sites whose URLs users ask the bot to read (see Privacy).
✅ supported · ⚙️ supported with an app-side replacement or extra configuration · ❌ not available (the model is told the capability is unavailable; there is never a silent fallback to OpenAI)
| Feature | OpenAI | Ollama | vLLM | OpenAI-compatible |
|---|---|---|---|---|
| Streaming replies | ✅ | ✅ | ✅ | ✅ |
| MetricsHub MCP tools | ✅ deferred namespaces + hosted tool search | ✅ plain function tools | ✅ plain function tools | ✅ plain function tools |
| Prometheus PromQL | ✅ | ✅ | ✅ | ✅ |
| MetricsHub config editing | ✅ | ✅ | ✅ | ✅ |
| Slack reactions / replies | ✅ | ✅ | ✅ | ✅ |
| Conversation state | ✅ server-side previous_response_id |
⚙️ app-side per-thread store | ⚙️ app-side per-thread store | ⚙️ app-side per-thread store |
| Context overflow handling | ✅ summarization with gpt-4o-mini |
⚙️ deterministic trimming to AI_CONTEXT_LENGTH |
⚙️ deterministic trimming | ⚙️ deterministic trimming |
| Context window detection | n/a | ✅ native API (/api/ps, /api/show) |
✅ max_model_len from /v1/models |
⚙️ context_length from /v1/models when reported, else set AI_CONTEXT_LENGTH |
| Knowledge base search | ✅ hosted file_search (vector stores) |
✅ local index, AI_EMBEDDING_MODEL defaults to nomic-embed-text |
⚙️ local index, needs a dedicated AI_EMBEDDING_BASE_URL |
⚙️ local index, needs AI_EMBEDDING_MODEL |
| Knowledge base writes | ✅ vector store upload | ✅ local Markdown + embeddings | ⚙️ same, once embeddings are configured | ⚙️ same, once embeddings are configured |
| Web search | ✅ hosted web_search_preview |
⚙️ app-side web_search (SearXNG or Ollama cloud) |
⚙️ app-side web_search (SearXNG or Ollama cloud) |
⚙️ app-side web_search (SearXNG or Ollama cloud) |
| Web page reading | ✅ done by hosted web search | ✅ fetch_url tool |
✅ fetch_url tool |
✅ fetch_url tool |
| Code execution | ✅ hosted Code Interpreter | ⚙️ local run_python (Pyodide) |
⚙️ local run_python (Pyodide) |
⚙️ local run_python (Pyodide) |
| Data-file attachments | ✅ OpenAI Files API | ⚙️ staged into the sandbox’s /data |
⚙️ staged into the sandbox’s /data |
⚙️ staged into the sandbox’s /data |
| Screenshots / images | ✅ native | ⚙️ described as text by OLLAMA_VISION_MODEL (unset = ❌) |
✅ native input_image via the media store |
⚙️ native with AI_IMAGE_INPUT=true + media store (default ❌) |
| PDF attachments | ✅ OpenAI Files API | ⚙️ staged into /data like any data file (no built-in PDF reader) |
⚙️ staged into /data |
⚙️ staged into /data |
| Reasoning summaries in logs | ✅ | ❌ | ❌ | ❌ |
| Model adopted when unset | fixed | default qwen3.8:27b |
✅ the single served model | ✅ when exactly one is served, else AI_MODEL required |
/v1/models at health check |
n/a | native API instead | required | optional (missing = warning, model unverified) |
Every self-hosted backend needs POST /v1/responses with stream: true and function tools.
A /v1/chat/completions-only server is not supported: there is no chat-completions mode, and
the bot does not translate between the two APIs. Ollama and vLLM ship a Responses endpoint; for a
gateway, check that it does before choosing openai-compatible. The startup health check sends one
real streaming /v1/responses request, which catches a chat-completions-only server; it does not
offer a function tool, so run the doctor script (whose probe does)
to confirm tool support before going live.
The only place the bot calls /v1/chat/completions is the Ollama vision sidecar
(OLLAMA_VISION_MODEL), because Ollama’s Responses endpoint has no image input.
Request fields sent per backend:
| Backend | Fields |
|---|---|
openai |
Full Responses API: previous_response_id, reasoning, text, hosted tools (file_search, web_search_preview, code_interpreter, tool search), safety_identifier |
ollama |
model, input, max_output_tokens, stream, plus tools when offered (stateless, no previous_response_id; the system prompt travels inside input) |
vllm |
Universal fields only, plus input_image items with the mandatory detail field; strict input conforming on |
openai-compatible |
Universal fields only (model, input, tools, stream, max_output_tokens); input_image and strict input are opt-in flags |
vllm is exactly openai-compatible with AI_IMAGE_INPUT=true, AI_STRICT_INPUT=true and a
mandatory /v1/models. Enable AI_STRICT_INPUT on a generic endpoint when the server answers
400 System message must be at the beginning or rejects assistant history carrying output_text
content.
Ollama’s /v1/responses is stateless: it implements neither previous_response_id nor
conversation. vLLM has a Responses store, but it is unbounded and process-local, so the bot does
not rely on it either. In all self-hosted modes the bot keeps history application-side, keyed by
team + channel + thread_ts, and resends the relevant user/assistant/tool items on every request.
History is trimmed deterministically to fit AI_CONTEXT_LENGTH, always retaining the system
prompt, the recent turns and intact tool-call/result pairs. The store is in-memory: after a
restart the text history is rebuilt from the Slack thread (the same cold-start path OpenAI mode
uses), losing only the tool-call detail of earlier turns.
In ollama, vllm and openai-compatible modes, prompts, documents and tool results never go to
OpenAI and never leave your network beyond your own inference endpoint, with two exceptions:
WEB_SEARCH_PROVIDER=ollama-cloud (opt-in) sends the search query (not the conversation) to
ollama.com.fetch_url (enabled by default; FETCH_URL_ENABLED=false removes it) contacts the sites whose URLs users paste (plus api.github.com for GitHub links)
with a plain GET carrying no Slack, MCP or GitHub credentials. Fetched pages are untrusted input;
the system prompt tells the model to treat them as data and never to follow instructions found in
them.In openai mode, conversation content, attachments and tool results are processed by OpenAI under
your OpenAI account’s data terms. Each request carries a privacy-preserving hash of the Slack
workspace and user IDs as safety_identifier.
| Backend | Model | Notes |
|---|---|---|
openai |
gpt-5.6-sol |
Reasoning effort medium, low verbosity, up to 8,000 output tokens |
ollama |
qwen3.8:27b + nomic-embed-text |
Optional qwen3-vl:8b-instruct-8k for screenshots; Ollama current release |
vllm |
Qwen 3.8 27B INT8 (multimodal) | vLLM 0.27+; embeddings from a second small instance or an Ollama server |
openai-compatible |
Any function-calling model | Set AI_CONTEXT_LENGTH explicitly; NVIDIA nv-embedqa embeddings get their input_type automatically |
M8B is developed against the first three. Other models work when they support function calling reliably; the doctor script’s function-tool probe is the fastest way to check one.