M8B Slack Bot

AI backends: support matrix

M8B talks to its model through the OpenAI Responses API (POST /v1/responses, streaming, function tools). AI_PROVIDER selects one of four backends. The application never branches on the backend name; it reads the provider’s capability flags, so the columns below are exactly what the code does. Variables are documented in CONFIGURATION.md.

AI_PROVIDER Backend Conversation data goes to
openai (default) Hosted OpenAI, gpt-5.6-sol OpenAI
ollama Ollama’s /v1/responses (text model, e.g. qwen3.8:27b) Your Ollama host only (search queries to ollama.com if WEB_SEARCH_PROVIDER=ollama-cloud)
vllm vLLM’s /v1/responses (multimodal model, e.g. Qwen 3.8 27B INT8) Your vLLM host only
openai-compatible Any other /v1/responses server: gateway, LiteLLM, NVIDIA NIM, proxy, … That endpoint, wherever it is

In every self-hosted mode the default-enabled page reader also contacts the public sites whose URLs users ask the bot to read (see Privacy).

Feature matrix

✅ supported · ⚙️ supported with an app-side replacement or extra configuration · ❌ not available (the model is told the capability is unavailable; there is never a silent fallback to OpenAI)

Feature OpenAI Ollama vLLM OpenAI-compatible
Streaming replies ✅ ✅ ✅ ✅
MetricsHub MCP tools ✅ deferred namespaces + hosted tool search ✅ plain function tools ✅ plain function tools ✅ plain function tools
Prometheus PromQL ✅ ✅ ✅ ✅
MetricsHub config editing ✅ ✅ ✅ ✅
Slack reactions / replies ✅ ✅ ✅ ✅
Conversation state ✅ server-side previous_response_id ⚙️ app-side per-thread store ⚙️ app-side per-thread store ⚙️ app-side per-thread store
Context overflow handling ✅ summarization with gpt-4o-mini ⚙️ deterministic trimming to AI_CONTEXT_LENGTH ⚙️ deterministic trimming ⚙️ deterministic trimming
Context window detection n/a ✅ native API (/api/ps, /api/show) ✅ max_model_len from /v1/models ⚙️ context_length from /v1/models when reported, else set AI_CONTEXT_LENGTH
Knowledge base search ✅ hosted file_search (vector stores) ✅ local index, AI_EMBEDDING_MODEL defaults to nomic-embed-text ⚙️ local index, needs a dedicated AI_EMBEDDING_BASE_URL ⚙️ local index, needs AI_EMBEDDING_MODEL
Knowledge base writes ✅ vector store upload ✅ local Markdown + embeddings ⚙️ same, once embeddings are configured ⚙️ same, once embeddings are configured
Web search ✅ hosted web_search_preview ⚙️ app-side web_search (SearXNG or Ollama cloud) ⚙️ app-side web_search (SearXNG or Ollama cloud) ⚙️ app-side web_search (SearXNG or Ollama cloud)
Web page reading ✅ done by hosted web search ✅ fetch_url tool ✅ fetch_url tool ✅ fetch_url tool
Code execution ✅ hosted Code Interpreter ⚙️ local run_python (Pyodide) ⚙️ local run_python (Pyodide) ⚙️ local run_python (Pyodide)
Data-file attachments ✅ OpenAI Files API ⚙️ staged into the sandbox’s /data ⚙️ staged into the sandbox’s /data ⚙️ staged into the sandbox’s /data
Screenshots / images ✅ native ⚙️ described as text by OLLAMA_VISION_MODEL (unset = ❌) ✅ native input_image via the media store ⚙️ native with AI_IMAGE_INPUT=true + media store (default ❌)
PDF attachments ✅ OpenAI Files API ⚙️ staged into /data like any data file (no built-in PDF reader) ⚙️ staged into /data ⚙️ staged into /data
Reasoning summaries in logs ✅ ❌ ❌ ❌
Model adopted when unset fixed default qwen3.8:27b ✅ the single served model ✅ when exactly one is served, else AI_MODEL required
/v1/models at health check n/a native API instead required optional (missing = warning, model unverified)

What each backend must provide

Every self-hosted backend needs POST /v1/responses with stream: true and function tools. A /v1/chat/completions-only server is not supported: there is no chat-completions mode, and the bot does not translate between the two APIs. Ollama and vLLM ship a Responses endpoint; for a gateway, check that it does before choosing openai-compatible. The startup health check sends one real streaming /v1/responses request, which catches a chat-completions-only server; it does not offer a function tool, so run the doctor script (whose probe does) to confirm tool support before going live.

The only place the bot calls /v1/chat/completions is the Ollama vision sidecar (OLLAMA_VISION_MODEL), because Ollama’s Responses endpoint has no image input.

Request fields sent per backend:

Backend Fields
openai Full Responses API: previous_response_id, reasoning, text, hosted tools (file_search, web_search_preview, code_interpreter, tool search), safety_identifier
ollama model, input, max_output_tokens, stream, plus tools when offered (stateless, no previous_response_id; the system prompt travels inside input)
vllm Universal fields only, plus input_image items with the mandatory detail field; strict input conforming on
openai-compatible Universal fields only (model, input, tools, stream, max_output_tokens); input_image and strict input are opt-in flags

vllm is exactly openai-compatible with AI_IMAGE_INPUT=true, AI_STRICT_INPUT=true and a mandatory /v1/models. Enable AI_STRICT_INPUT on a generic endpoint when the server answers 400 System message must be at the beginning or rejects assistant history carrying output_text content.

Conversation state in self-hosted modes

Ollama’s /v1/responses is stateless: it implements neither previous_response_id nor conversation. vLLM has a Responses store, but it is unbounded and process-local, so the bot does not rely on it either. In all self-hosted modes the bot keeps history application-side, keyed by team + channel + thread_ts, and resends the relevant user/assistant/tool items on every request. History is trimmed deterministically to fit AI_CONTEXT_LENGTH, always retaining the system prompt, the recent turns and intact tool-call/result pairs. The store is in-memory: after a restart the text history is rebuilt from the Slack thread (the same cold-start path OpenAI mode uses), losing only the tool-call detail of earlier turns.

Privacy

In ollama, vllm and openai-compatible modes, prompts, documents and tool results never go to OpenAI and never leave your network beyond your own inference endpoint, with two exceptions:

In openai mode, conversation content, attachments and tool results are processed by OpenAI under your OpenAI account’s data terms. Each request carries a privacy-preserving hash of the Slack workspace and user IDs as safety_identifier.

Reference models

Backend Model Notes
openai gpt-5.6-sol Reasoning effort medium, low verbosity, up to 8,000 output tokens
ollama qwen3.8:27b + nomic-embed-text Optional qwen3-vl:8b-instruct-8k for screenshots; Ollama current release
vllm Qwen 3.8 27B INT8 (multimodal) vLLM 0.27+; embeddings from a second small instance or an Ollama server
openai-compatible Any function-calling model Set AI_CONTEXT_LENGTH explicitly; NVIDIA nv-embedqa embeddings get their input_type automatically

M8B is developed against the first three. Other models work when they support function calling reliably; the doctor script’s function-tool probe is the fastest way to check one.