BrainAPI
BrainAPI
Start

For the complete BrainAPI documentation index, see llms.txt. A markdown version of any docs page is available by appending .md to its URL. Docs MCP: /docs/mcp.

Installation

Install BrainAPI v2 for development or production

For agents

  • Recommended: npm install -g [email protected] → brainapi init → brainapi start
  • The TUI manages source, a Python venv, .env, services, API, worker, MCP, and optional plugins
  • Contributors can instead clone Lumen-Labs/brainapi2, copy .env.example → .env, and run from source
  • Dev: ./bin/brainapi <command> · Prod: docker compose -f example-docker-compose.yaml up -d
  • API default :8000; product MCP often :8001; docs MCP is separate at /docs/mcp
  • After plugin install, restart API (and worker/MCP when those surfaces change)

BrainAPI v2 supports three installation paths. Use the TUI for a local evaluation or application-development environment, clone the repository when contributing to BrainAPI itself, and use an explicit Compose deployment for a controlled self-hosted environment.

The brainapi TUI is the shortest supported path from a new machine to a running API, worker, MCP server, console, and selected backing services.

npm install -g [email protected]
brainapi init
brainapi start

brainapi init checks Python and Docker, clones BrainAPI into ~/.brainapi/source, creates its virtual environment, and opens the configuration wizard. The wizard lets you choose databases, local or remote model providers, ingestion mode, Search, credentials, services, and official plugins. It writes the resulting configuration to ~/.brainapi/source/.env.

When brainapi start is ready, the common local endpoints are:

SurfaceURL
REST APIhttp://localhost:8000
Web consolehttp://localhost:8000/console
Product MCP serverhttp://localhost:8001/mcp

Verify the API with the system PAT created or supplied during setup:

export BRAINPAT_TOKEN="replace-with-your-token"
curl --fail http://localhost:8000/ \
  -H "BrainPAT: $BRAINPAT_TOKEN"

Expected output is ok. If startup or verification fails, run brainapi doctor; it checks the managed Python environment, Docker, model providers, credentials, and configured services.

brainapi doctor

Use brainapi config to revisit the wizard and brainapi update to update the managed checkout. See the complete TUI installer reference for non-interactive flags, Search choices, service selection, directory layout, and process controls.

Production stack overview

The provided version: "3.8" compose stack includes:

  • nginx as reverse proxy and TLS termination
  • redis for cache and task queue backend
  • mongo for document storage
  • neo4j for graph relationships
  • etcd and minio as Milvus dependencies
  • milvus as vector database
  • brainapi as API gateway (:8000)
  • brainapi-worker as Celery worker
  • brainapi-mcp as MCP server (:8001)

The API, worker, and MCP containers share the same image and rely on the same environment variables and plugin volume. Keep these values aligned across all three services.

Production setup

1) Prepare host paths referenced by compose

mkdir -p /srv/nginx/conf/conf.d
mkdir -p /srv/nginx/logs
mkdir -p /root

The compose file expects these host mounts:

  • /srv/nginx/conf/nginx.conf
  • /srv/nginx/conf/conf.d
  • /etc/letsencrypt
  • /srv/nginx/logs
  • /root/.env
  • /root/gcp_credentials.json

2) Configure environment variables

Copy your environment file to /root/.env and set production values.

cp .env.example /root/.env

At minimum, verify:

Prop

Type

3) Review service credentials before startup

The example compose file ships with development defaults. Update these values before production usage:

  • NEO4J_AUTH=neo4j/your_password
  • MONGO_INITDB_ROOT_USERNAME
  • MONGO_INITDB_ROOT_PASSWORD
  • MINIO_ACCESS_KEY
  • MINIO_SECRET_KEY

4) Start the stack

docker compose -f example-docker-compose.yaml pull
docker compose -f example-docker-compose.yaml up -d
docker compose -f example-docker-compose.yaml ps

5) Verify endpoints

curl -f http://localhost:8000/ -H "BrainPAT: $BRAINPAT_TOKEN"
curl -f http://localhost:8001/
curl -f http://localhost:9091/healthz
curl -f http://localhost:9000/minio/health/live

If all checks pass, BrainAPI API, worker, and MCP services are running with Redis, MongoDB, Neo4j, and Milvus.

LLM providers

BrainAPI selects chat and embedding adapters via LLM_SMALL_PROVIDER, LLM_LARGE_PROVIDER, and EMBEDDINGS_PROVIDER (see .env.example).

DeepSeek (chat only)

DeepSeek is an OpenAI-compatible chat provider. It does not serve embeddings — keep EMBEDDINGS_PROVIDER on openai, azure, gcp_vertex, or another embedding-capable adapter.

LLM_SMALL_PROVIDER="deepseek"
LLM_LARGE_PROVIDER="deepseek"
DEEPSEEK_API_KEY="sk-..."
DEEPSEEK_SMALL_LLM_MODEL="deepseek-v4-flash"
DEEPSEEK_LARGE_LLM_MODEL="deepseek-v4-pro"
# Embeddings stay elsewhere, for example:
# EMBEDDINGS_PROVIDER="openai"

The TUI setup wizard can also select DeepSeek as a remote chat provider when configuring models.

Pipeline and ingest tuning

Defaults below match .env.example. Accurate pipeline mode no longer implies Observations or graph consolidation are always on.

Prop

Type

Advanced Architect knobs

These are optional cost/latency controls (documented in .env.example):

  • INGEST_ARCHITECT_PER_UNIT (default true) — Architect sees Scout chunks plus a prior window
  • INGEST_ARCHITECT_MAX_SCHEMA_CALLS (default 3)
  • INGEST_ARCHITECT_ESCALATE / INGEST_ARCHITECT_ESCALATE_MAX_TURNS
  • INGEST_ARCHITECT_DENSE_ENTITY_THRESHOLD / INGEST_ARCHITECT_DENSE_MAX_CHARS
  • INGEST_ARCHITECT_PRIOR_CONTEXT (auto | scratchpad | raw)
  • INGEST_ARCHITECT_SCRATCHPAD_TOKEN_CAP

Stage latency profiler

For /retrieve/context, set profile_stages: true on the request to receive stage_timings (see Context). Operators can also set TRACE_STAGE_PROFILER_ENABLED on the server (not shipped in .env.example).

Search configuration

Ranked Search is off by default and requires PostgreSQL for text chunks:

DATA_DB="postgresql"
SEARCH_ENABLED="true"
SEARCH_USE_DENSE="true"
SEARCH_USE_BM25="true"
SEARCH_FUSION="rrf"
SEARCH_FUSION_ALPHA="0.5"
SEARCH_BM25_K1="1.2"
SEARCH_BM25_B="0.75"
SEARCH_COMMUNITY_LABELS="TYPE,CLASS,TOPIC"
SEARCH_NEIGHBOR_FANOUT="50"
CONTEXT_PASSAGE_MODE="hybrid"

Prop

Type

Search startup validation rejects SEARCH_ENABLED=true unless DATA_DB=postgresql and at least one of SEARCH_USE_DENSE or SEARCH_USE_BM25 is true.

English full-text indexing is always available. To add another tokenizer for selected brains, set both values:

SEARCH_FTS_REGCONFIG="italian"
SEARCH_FTS_BRAINS="catalogit,catalogit2"

Supported alternate configurations are italian, spanish, and simple. Only listed brains use the alternate tsvector; other brains retain English.

Advanced candidate fill

SEARCH_LITERAL_FILL=true enables an opt-in literal-overlap sidecar. It scans token-specific chunk candidates and merges residual literal matches after the frozen core head. Keep it false unless the workload has been evaluated: it can add data reads and is not part of the default Search claim.

Large embedding dimensions

pgvector HNSW's regular vector index is limited for high dimensions. With Search enabled and an embedding dimension above 2000, BrainAPI creates a halfvec cosine HNSW index for candidate generation, then orders the candidate set with float32 vector distances. If the halfvec index cannot be created, startup fails rather than silently shipping an unindexed Search path.

The default Search latency target is p50 <200 ms excluding the separately profiled embed.query stage. It applies to mode=default, no plugin reranker, and no target; deeper catalog, plugin, and personalized requests have their own cost profile.

Next steps

Edit on GitHub

Last updated on

On this page