Enterprise brain

The Enterprise brain gives every Spark agent, and every AI client you connect over MCP, one shared understanding of your company data: which databases exist, how their tables relate, which procedures answered a question last time, and which company tools an agent may call. It lives in your Core Hub, on your infrastructure, under your own rules.

Why it exists

Executives ask questions faster than BI teams can publish views. A spreadsheet estate grows because each answer needs its own extract. The Enterprise brain starts where Gluesync is already strong, integrating databases, and turns the schema Gluesync already knows into something an agent can reason over. From there it extends through skills (reusable procedures) and company tools (approved remote MCP servers such as GitLab or Notion) you plug in while in platform.

Because the brain is exposed through the Core Hub MCP server, the same knowledge is reachable from Claude, ChatGPT, Grok, Meta Muse, Cursor, Hermes Agent, OpenClaw, or any other MCP-capable client. They ask; Core Hub answers with the caller’s permissions.

What the brain stores

The brain lives in a dedicated SQLite database, gluesync-ai-brain.db, beside the existing gluesync.db in the Core Hub data directory. It is included in support snapshots and backups. It contains:

Store Content

Schema graph

One node per schema, table, and column discovered through the SQL-capable agents of every pipeline, with column names, types, and key roles. When a connector reports a foreign-key target as schema.table on the column, the walk also stores that relationship as an edge (PostgreSQL, CockroachDB, YugabyteDB, Redshift, Vertica, MySQL, MariaDB, Oracle, SQL Server, and Db2 LUW). GridGain SQL has no foreign keys, so its columns stay as nodes without edges. Nodes are grouped by build generation so a rebuild swaps in atomically and never shows a half-built graph.

Agent memory

Short procedural notes an agent writes when a run or a chat turn completes: which tables answered which question, which filter mattered, which skill worked. Memory records location and procedure, never result rows and never secrets. Password, token, secret, and API key assignments are redacted before anything is stored.

Conversations

AI Studio conversations. They are owner-private: a user reads only their own conversations, and only SUPER_ADMIN keeps the existing cross-owner visibility rule.

Skill attachments

Published skills attached to a published agent without cutting a new agent version.

Approved remote tools

Registered HTTPS MCP servers (company tools), their host allow-list, an optional encrypted access token, and the last discovered tool snapshot. Core Hub never returns the token after save; the API only reports whether one is stored.

Installations that ran an earlier 2.3 preview keep their old AI_STUDIO_CONVERSATIONS rows inside gluesync.db. They are neither migrated nor deleted; new conversations are written only to the brain database.

Shared by default, private on request

Memory is partitioned by the published agent slug. Everyone who can chat with an agent benefits from what that agent has already learned.

  • A memory is SHARED unless the user asked otherwise. Phrases such as "keep this private" or "private memory" in the request mark the resulting memory PRIVATE to that user.

  • Private memories are visible only to their owner. Forgetting a memory (DELETE /api/ai/v1/agents/{slug}/memories/{memoryId}) works on shared memories and on your own private memories, never on somebody else’s.

  • Memory text is capped at 2,000 characters and carries entity hints (table names it mentions) so keyword search can find it.

Build the brain

  1. Open AI Studio → Enterprise brain.

  2. Follow the setup wizard (intro → search mode → build). Choose a search mode (see below) and start the catalog walk.

  3. When the brain is ready, the same workspace shows status, search mode, company tools, and the Connect your AI client panel.

Core Hub answers immediately with status CREATING and the message Creating Enterprise brain…​. The UI polls and shows tables and foreign-key edges seen so far. When the walk finishes, the status becomes READY (Enterprise brain is ready.); an interrupted walk resumes from its last checkpoint (pipeline, agent, schema, table) the next time you start it, so a large estate never restarts from zero. You can Run setup again or Rebuild graph from the dashboard without leaving the tab.

The walk streams schemas, tables, and columns page by page through the same discovery calls Query Studio uses, off the request thread, and never loads a whole catalog into memory. Query Studio’s own 40-table prompt cap is unchanged: the brain returns a subgraph for a question, not a prompt dump.

Only a Manager or Super admin (or a platform API key with the skills:publish scope) can change search settings or start a build. Everyone with AI Studio access can read the status.

Keyword or embeddings

Search mode is a user-visible setting, not a hidden add-on.

Mode Behaviour

Keyword (always available)

Full-text search (SQLite FTS5) over table and column names plus a foreign-key walk from every hit. Works with any chat model and needs no extra provider call.

Embeddings (suggested when available)

Offered as the suggested choice as soon as a catalog model with the EMBEDDINGS capability exists (for example an OpenAI-compatible embedding model). Table nodes and memories are embedded asynchronously and blended with the keyword results by cosine similarity, so a question that shares no words with a table name can still find it. Unavailable, and shown as such, until an embedding model is configured.

The embedding model is pinned with the brain. Changing the pinned model clears stored vectors and rebuilds them; switching back to keyword stops embedding calls altogether. The agent’s chat model is independent of this choice.

How agents use it

Before each chat turn or run step, Core Hub asks the brain for a preface: up to eight relevant memories visible to the caller and up to eight schema-graph lines for the question. The preface is prepended to the agent instructions. When a run or conversation reaches a terminal state, the input and the tool trace are sedimented into a memory candidate and queued; a background worker drains the queue so the request path never waits on ingestion.

Agents that write skills

A published agent can call the publish_skill capability to turn a procedure it just executed into an immutable INSTRUCTION skill. The chat transcript shows each produced skill as a link (Skill <slug> v`<version>`) that opens and highlights it in Build & publish. Runs list the same references under skillRefs.

publish_skill requires the caller to be a Manager or Super admin, or a platform key carrying the skills:publish scope. Attaching an existing skill to an agent (POST /api/ai/v1/agents/{slug}/skills with { "skillId": "slug:version" }) uses the same rule and does not create a new agent version.

Approved company tools (remote MCP)

Under Company tools on the Enterprise brain dashboard, a publisher selects Add company tool:

  • A gallery lists dozens of vendor-hosted MCP endpoints (for example Slack, Gmail, GitHub, GitLab.com, Notion, Linear, Atlassian, HubSpot, Stripe, Sentry, and many others). Pick a tile to pre-fill the HTTPS URL, display name, and allow-listed host. Some entries need one tenant-specific segment (organization, shop subdomain, or self-managed host).

  • If the tool you need is missing, use Don’t find the tool you’re looking for? → Add a custom MCP server and fill the same connection form by hand.

  • The URL must be https://, and its host must appear in the Allow-listed hosts field. Any call to a host outside the allow-list is refused.

  • When the vendor needs credentials, paste an Access token (or API key). Core Hub stores it encrypted with the platform storage key, sends it as a bearer token on discovery and tool calls, and never returns the value to the UI. Public servers leave the token empty. The dashboard shows Token stored when a credential is present.

  • Register & discover snapshots the server’s tools/list result. Select the tools you want, give the bundle a slug and instructions, and Publish to create a TOOL_BUNDLE skill whose capabilities are named remote:<serverId>:<toolName>.

Agents that carry the bundle can call those tools during a run. Remote calls are treated as writes: they wait for a signed-in approval before executing, and runs started by an API key execute as the key owner.

Reach the brain from any AI client

The Connect your AI client panel on the Enterprise brain tab copies a setup prompt for Claude, ChatGPT, Cursor, and similar MCP clients. It includes the Core Hub MCP endpoint and points you at a Personal API Token under Settings → Account → Personal API Tokens.

Cloud-hosted connectors (for example Claude or ChatGPT cloud MCP) cannot reach a Core Hub that is only on localhost or a private address. The panel warns when the Control Plane origin is local-only and the setup prompt tells the client to use a local stdio bridge (such as mcp-remote) or a public HTTPS URL. Local agents that can open loopback (Cursor on the same machine) do not need that bridge.

The Core Hub MCP server exposes an AI agents category so external clients can use published agents and the brain without the Control Plane:

Tool Purpose

list_published_agents

Published agent slugs and latest versions the caller may invoke.

start_agent_run

Start a durable run of a published agent with a prompt; returns the run id.

get_agent_run

Read a run’s status, output, and skillRefs.

publish_skill

Publish an instruction skill (skills:publish or Manager/Super admin).

search_enterprise_brain

Search the schema graph and the caller’s visible memories. Returns tables, columns, and procedures, never row data.

Configure the client as described in Connecting a client; every call runs under the caller’s token, so the brain never reveals a private memory or a table the caller cannot see.

REST API

All routes are under /api/ai/v1 and accept a signed-in bearer token or a gsa_ platform key.

GET    /brain/settings                       search mode, pinned model, suggested model
PUT    /brain/settings                       {"searchMode":"KEYWORD"|"EMBEDDINGS","embeddingModelId":...}
GET    /brain/job                            CREATING | READY | FAILED | IDLE with table and edge counts
POST   /brain/job                            start or resume the catalog walk
GET    /brain/search?query=&agentSlug=       subgraph lines and visible memory text
GET    /agents/{slug}/memories               shared memories plus the caller's private memories
DELETE /agents/{slug}/memories/{memoryId}    forget
POST   /agents/{slug}/memories/{memoryId}/private
POST   /agents/{slug}/skills                 attach a published skill
GET    /remote-mcp                           approved servers (`hasSecret`, never the token)
POST   /remote-mcp                           register an HTTPS server; optional write-only `secret`
POST   /remote-mcp/{id}/discover             snapshot tools (uses the stored encrypted token)
POST   /remote-mcp/{id}/skills               publish a TOOL_BUNDLE skill
DELETE /remote-mcp/{id}

POST /remote-mcp accepts a write-only secret (bearer token). Core Hub encrypts it at rest. Omitting secret on a later save keeps the existing credential so a rediscovery does not drop it. List responses expose hasSecret only.

The authenticated OpenAPI contract at /api/ai/v1/openapi.yaml documents request and response bodies.

Observability and limits

  • Spans and metrics carry identifiers and counts only, never memory text, prompts, or table contents.

  • Memory text is capped at 2,000 characters; table node text at 8,000 characters; up to 400 stored vectors per model are ranked per query.

  • Ingestion is a queue (AI_BRAIN_INGESTION) drained in batches of 32 by a Core Hub background worker.