# Installation (https://brainapi.lumen-labs.ai/docs/v2/installation)

> For the complete BrainAPI documentation index, see [llms.txt](https://brainapi.lumen-labs.ai/docs/llms.txt). A markdown version of any docs page is available by appending `.md` to its URL (e.g. https://brainapi.lumen-labs.ai/docs/v2/installation.md).

Install BrainAPI v2 for development or production

<AgentNote>
- Recommended: `npm install -g brainapi-tui@0.4.0` → `brainapi init` → `brainapi start`
- The TUI manages source, a Python venv, `.env`, services, API, worker, MCP, and optional plugins
- Contributors can instead clone `Lumen-Labs/brainapi2`, copy `.env.example` → `.env`, and run from source
- Dev: `./bin/brainapi <command>` · Prod: `docker compose -f example-docker-compose.yaml up -d`
- API default `:8000`; product MCP often `:8001`; docs MCP is separate at `/docs/mcp`
- After plugin install, restart API (and worker/MCP when those surfaces change)
</AgentNote>

BrainAPI v2 supports three installation paths. Use the TUI for a local evaluation or application-development environment, clone the repository when contributing to BrainAPI itself, and use an explicit Compose deployment for a controlled self-hosted environment.

<Tabs items={["TUI (recommended)", "Development", "Production (Docker Compose)"]}>
  <Tab value="TUI (recommended)">
  The `brainapi` TUI is the shortest supported path from a new machine to a running API, worker, MCP server, console, and selected backing services.

```bash
npm install -g brainapi-tui@0.4.0
brainapi init
brainapi start
```

`brainapi init` checks Python and Docker, clones BrainAPI into `~/.brainapi/source`, creates its virtual environment, and opens the configuration wizard. The wizard lets you choose databases, local or remote model providers, ingestion mode, Search, credentials, services, and official plugins. It writes the resulting configuration to `~/.brainapi/source/.env`.

When `brainapi start` is ready, the common local endpoints are:

| Surface | URL |
| --- | --- |
| REST API | `http://localhost:8000` |
| Web console | `http://localhost:8000/console` |
| Product MCP server | `http://localhost:8001/mcp` |

Verify the API with the system PAT created or supplied during setup:

```bash
export BRAINPAT_TOKEN="replace-with-your-token"
curl --fail http://localhost:8000/ \
  -H "BrainPAT: $BRAINPAT_TOKEN"
```

Expected output is `ok`. If startup or verification fails, run `brainapi doctor`; it checks the managed Python environment, Docker, model providers, credentials, and configured services.

```bash
brainapi doctor
```

Use `brainapi config` to revisit the wizard and `brainapi update` to update the managed checkout. See the complete [TUI installer reference](https://brainapi.lumen-labs.ai/docs/v2/tui) for non-interactive flags, Search choices, service selection, directory layout, and process controls.
  </Tab>
  <Tab value="Development">
  Use this flow if you are working on the BrainAPI codebase or building plugins.

```bash
git clone git@github.com:Lumen-Labs/brainapi2.git
cd brainapi2
cp .env.example .env
chmod +x bin/brainapi.sh
./bin/brainapi <command> [options]
```

Fill the values in `.env` before running API, worker, MCP, or plugin workflows.
  </Tab>
  <Tab value="Production (Docker Compose)">
  Use this flow for self-hosted environments. The `example-docker-compose.yaml` file starts the full stack used by BrainAPI v2.
  Compose file: [https://github.com/Lumen-Labs/brainapi2/blob/main/example-docker-compose.yaml](https://github.com/Lumen-Labs/brainapi2/blob/main/example-docker-compose.yaml)
  </Tab>
</Tabs>

## Production stack overview

The provided `version: "3.8"` compose stack includes:

- `nginx` as reverse proxy and TLS termination
- `redis` for cache and task queue backend
- `mongo` for document storage
- `neo4j` for graph relationships
- `etcd` and `minio` as Milvus dependencies
- `milvus` as vector database
- `brainapi` as API gateway (`:8000`)
- `brainapi-worker` as Celery worker
- `brainapi-mcp` as MCP server (`:8001`)

<Callout type="info">
  The API, worker, and MCP containers share the same image and rely on the same
  environment variables and plugin volume. Keep these values aligned across all
  three services.
</Callout>

## Production setup

### 1) Prepare host paths referenced by compose

```bash
mkdir -p /srv/nginx/conf/conf.d
mkdir -p /srv/nginx/logs
mkdir -p /root
```

The compose file expects these host mounts:

- `/srv/nginx/conf/nginx.conf`
- `/srv/nginx/conf/conf.d`
- `/etc/letsencrypt`
- `/srv/nginx/logs`
- `/root/.env`
- `/root/gcp_credentials.json`

### 2) Configure environment variables

Copy your environment file to `/root/.env` and set production values.

```bash
cp .env.example /root/.env
```

At minimum, verify:

<TypeTable
  type={{
    BRAINPAT_TOKEN: {
      description: "Token used by protected health checks and authenticated API calls.",
      type: "string",
      required: true,
    },
    BRAINAPI_PLUGINS: {
      description: "Comma-separated plugin list loaded at startup.",
      type: "string",
      required: false,
      default: "my-plugin:1.0.0,analytics-plugin",
    },
    PLUGIN_REGISTRY_URL: {
      description: "Custom plugin registry URL if not using the default.",
      type: "string",
      required: false,
    },
    PLUGIN_PUBLISHER_ID: {
      description: "Publisher identifier for registry operations.",
      type: "string",
      required: false,
    },
    PLUGIN_PUBLISHER_API_KEY: {
      description: "Publisher API key for registry operations.",
      type: "string",
      required: false,
    },
  }}
/>

### 3) Review service credentials before startup

The example compose file ships with development defaults. Update these values before production usage:

- `NEO4J_AUTH=neo4j/your_password`
- `MONGO_INITDB_ROOT_USERNAME`
- `MONGO_INITDB_ROOT_PASSWORD`
- `MINIO_ACCESS_KEY`
- `MINIO_SECRET_KEY`

### 4) Start the stack

```bash
docker compose -f example-docker-compose.yaml pull
docker compose -f example-docker-compose.yaml up -d
docker compose -f example-docker-compose.yaml ps
```

### 5) Verify endpoints

```bash
curl -f http://localhost:8000/ -H "BrainPAT: $BRAINPAT_TOKEN"
curl -f http://localhost:8001/
curl -f http://localhost:9091/healthz
curl -f http://localhost:9000/minio/health/live
```

If all checks pass, BrainAPI API, worker, and MCP services are running with Redis, MongoDB, Neo4j, and Milvus.

## LLM providers

BrainAPI selects chat and embedding adapters via `LLM_SMALL_PROVIDER`, `LLM_LARGE_PROVIDER`, and `EMBEDDINGS_PROVIDER` (see `.env.example`).

### DeepSeek (chat only)

DeepSeek is an OpenAI-compatible chat provider. It does **not** serve embeddings — keep `EMBEDDINGS_PROVIDER` on `openai`, `azure`, `gcp_vertex`, or another embedding-capable adapter.

```bash
LLM_SMALL_PROVIDER="deepseek"
LLM_LARGE_PROVIDER="deepseek"
DEEPSEEK_API_KEY="sk-..."
DEEPSEEK_SMALL_LLM_MODEL="deepseek-v4-flash"
DEEPSEEK_LARGE_LLM_MODEL="deepseek-v4-pro"
# Embeddings stay elsewhere, for example:
# EMBEDDINGS_PROVIDER="openai"
```

The [TUI setup wizard](https://brainapi.lumen-labs.ai/docs/v2/tui) can also select DeepSeek as a remote chat provider when configuring models.

## Pipeline and ingest tuning

Defaults below match `.env.example`. Accurate pipeline mode no longer implies Observations or graph consolidation are always on.

<TypeTable
  type={{
    PIPELINE_MODE: {
      description: "Ingestion strength: accurate (full agents) or lightweight.",
      type: "string",
      required: false,
      default: "accurate",
    },
    RUN_OBSERVATIONS: {
      description: "When true, ObservationsAgent writes LLM notes (does not mutate the graph).",
      type: "boolean",
      required: false,
      default: "false",
    },
    RUN_GRAPH_CONSOLIDATOR: {
      description: "When true, consolidates the graph after ingest (LLM mutations). Keep false until audited.",
      type: "boolean",
      required: false,
      default: "false",
    },
    INGEST_ARCHITECT_MODE: {
      description: "Architect extract path: batch (default), schema, or tooler.",
      type: "string",
      required: false,
      default: "batch",
    },
    INGEST_DEFER_JANITOR: {
      description: "Batched post-Architect Janitor when true; per-create Janitor when false.",
      type: "boolean",
      required: false,
      default: "true",
    },
    INGEST_JANITOR_MAX_LLM_CALLS: {
      description: "Hard cap on Janitor LLM batches per session; remainder dropped with audit.",
      type: "number",
      required: false,
      default: "2",
    },
    JANITOR_BATCH_SIZE: {
      description: "Relationships per Janitor LLM call when batching.",
      type: "number",
      required: false,
      default: "20",
    },
  }}
/>

### Advanced Architect knobs

These are optional cost/latency controls (documented in `.env.example`):

- `INGEST_ARCHITECT_PER_UNIT` (default `true`) — Architect sees Scout chunks plus a prior window
- `INGEST_ARCHITECT_MAX_SCHEMA_CALLS` (default `3`)
- `INGEST_ARCHITECT_ESCALATE` / `INGEST_ARCHITECT_ESCALATE_MAX_TURNS`
- `INGEST_ARCHITECT_DENSE_ENTITY_THRESHOLD` / `INGEST_ARCHITECT_DENSE_MAX_CHARS`
- `INGEST_ARCHITECT_PRIOR_CONTEXT` (`auto` | `scratchpad` | `raw`)
- `INGEST_ARCHITECT_SCRATCHPAD_TOKEN_CAP`

### Stage latency profiler

For `/retrieve/context`, set `profile_stages: true` on the request to receive `stage_timings` (see [Context](https://brainapi.lumen-labs.ai/docs/v2/retrieval/context#stage-profiler)). Operators can also set `TRACE_STAGE_PROFILER_ENABLED` on the server (not shipped in `.env.example`).

## Search configuration

Ranked Search is off by default and requires PostgreSQL for text chunks:

```dotenv
DATA_DB="postgresql"
SEARCH_ENABLED="true"
SEARCH_USE_DENSE="true"
SEARCH_USE_BM25="true"
SEARCH_FUSION="rrf"
SEARCH_FUSION_ALPHA="0.5"
SEARCH_BM25_K1="1.2"
SEARCH_BM25_B="0.75"
SEARCH_COMMUNITY_LABELS="TYPE,CLASS,TOPIC"
SEARCH_NEIGHBOR_FANOUT="50"
CONTEXT_PASSAGE_MODE="hybrid"
```

<TypeTable
  type={{
    SEARCH_ENABLED: {
      description: "Enable GET|POST /retrieve/search and PostgreSQL search indexes.",
      type: "boolean",
      required: false,
      default: "false",
    },
    SEARCH_USE_DENSE: {
      description: "Use embedding-vector retrieval for the passages channel.",
      type: "boolean",
      required: false,
      default: "true",
    },
    SEARCH_USE_BM25: {
      description: "Use PostgreSQL BM25 retrieval for the passages channel.",
      type: "boolean",
      required: false,
      default: "true",
    },
    SEARCH_FUSION: {
      description: "Default passage fusion: rrf or cc.",
      type: "string",
      required: false,
      default: "rrf",
    },
    SEARCH_FUSION_ALPHA: {
      description: "Dense weight used by convex-combination fusion.",
      type: "number",
      required: false,
      default: "0.5",
    },
    SEARCH_BM25_K1: {
      description: "Positive BM25 term-frequency saturation constant.",
      type: "number",
      required: false,
      default: "1.2",
    },
    SEARCH_BM25_B: {
      description: "BM25 document-length normalization, from 0 through 1.",
      type: "number",
      required: false,
      default: "0.75",
    },
    SEARCH_COMMUNITY_LABELS: {
      description: "Comma-separated graph hub labels used by the communities channel.",
      type: "string",
      required: false,
      default: "TYPE,CLASS,TOPIC",
    },
    SEARCH_NEIGHBOR_FANOUT: {
      description: "Maximum one-hop members inspected per seed for neighbor/community expansion.",
      type: "number",
      required: false,
      default: "50",
    },
    CONTEXT_PASSAGE_MODE: {
      description: "Passage leg for /retrieve/context while Search is enabled: hybrid, bm25, dense, or ilike.",
      type: "string",
      required: false,
      default: "hybrid",
    },
  }}
/>

Search startup validation rejects `SEARCH_ENABLED=true` unless
`DATA_DB=postgresql` and at least one of `SEARCH_USE_DENSE` or
`SEARCH_USE_BM25` is true.

### Multilingual full-text search

English full-text indexing is always available. To add another tokenizer for
selected brains, set both values:

```dotenv
SEARCH_FTS_REGCONFIG="italian"
SEARCH_FTS_BRAINS="catalogit,catalogit2"
```

Supported alternate configurations are `italian`, `spanish`, and `simple`.
Only listed brains use the alternate `tsvector`; other brains retain English.

### Advanced candidate fill

`SEARCH_LITERAL_FILL=true` enables an opt-in literal-overlap sidecar. It scans
token-specific chunk candidates and merges residual literal matches after the
frozen core head. Keep it false unless the workload has been evaluated: it can
add data reads and is not part of the default Search claim.

### Large embedding dimensions

pgvector HNSW's regular vector index is limited for high dimensions. With
Search enabled and an embedding dimension above 2000, BrainAPI creates a
`halfvec` cosine HNSW index for candidate generation, then orders the candidate
set with float32 vector distances. If the halfvec index cannot be created,
startup fails rather than silently shipping an unindexed Search path.

The default Search latency target is p50 `<200 ms` excluding the separately
profiled `embed.query` stage. It applies to `mode=default`, no plugin reranker,
and no `target`; deeper catalog, plugin, and personalized requests have their
own cost profile.

## Next steps

- [Structured ingestion](https://brainapi.lumen-labs.ai/docs/v2/ingestion/structured-data)
- [Retrieve context](https://brainapi.lumen-labs.ai/docs/v2/retrieval/context)
- [Ranked Search](https://brainapi.lumen-labs.ai/docs/v2/retrieval/search)
