Nacre

Connectors

Six sources, kept in sync. Nothing crawls.

Nacre takes documents pushed in over REST, MCP, the SDK or the command line, and you own the sync from your systems. A connector is that sync written once, for a source many teams have: one container per source, configured entirely through its environment, that adds what the source has, changes what it changed, and removes what it dropped. Apache 2.0, like the core.

git
The text files of a git repository A path rule picks the layer. A blob's hash is the version, so a sweep reads only what moved, and a path gone from the tree leaves the index.
s3
The objects of a bucket — AWS, MinIO, any S3 store A key rule picks the layer and the ETag is the version. A PDF or a Word file goes up as the file, and the index extracts it.
sql
The rows a query returns — Postgres, MySQL, MariaDB A template over the columns picks the layer. A watermark column is the version where the table has one; otherwise the row's own hash is.
drive
A Google Drive folder, or a shared drive Files go up as files, and a Google Doc as the Word file it exports to. A file that was trashed or deleted leaves the index.
mongo
The documents a filter matches in a MongoDB collection A template over the fields picks the layer, embedded fields included. A document edited out of the filter is a document that leaves the index.
imap
The messages of a mailbox folder A template over the headers picks the layer. An attachment the index reads is a document of its own; a corrected draft is a change, not a second message.

How a sweep works

Add and change are easy. Remove is the whole difficulty.

Ingest is idempotent on the layer and the document's own id, so adding and changing cost a connector nothing but sending what it sees. The index cannot know that a source dropped something — that is what each connector's state is for.

Removal needs a complete listing

A document is removed when a complete listing of the source no longer contains it. A listing that breaks off halfway removes nothing: the half it never reached has not been called gone, and that is the first thing the suite checks.

The version is the cheap path

Where a source offers a version — a blob hash, an ETag, a watermark, a UID — an unchanged item is skipped without being fetched. A repository of ten thousand files costs a sweep the ones that moved.

Write, and only write

A connector holds a service account key with write on the layers it maps to, and nothing else. It never searches, and in Nacre write does not imply read — so the key that feeds an index cannot read it back.


One container per source

Run one, then watch it.

Every connector reads the same six variables, plus its own for the source, and refuses to start on a missing or malformed one rather than syncing nothing and reporting success.

docker run

A repository into two layers

Every document carries metadata.connector and metadata.source, so a search can be narrowed to what one connector brought.

$ docker run -v git-state:/state \
    -e NACRE_URL=https://nacre.example.com \
    -e NACRE_TOKEN=… \
    -e GIT_URL=https://github.com/acme/handbook \
    -e 'GIT_LAYERS=docs/**=handbook;src/**=code' \
    ghcr.io/nacre-work/connectors/git:0.1.1

Every connector

Six variables, and a port

/healthz, /status and /metrics answer on the port. /status is a versioned contract, and the metrics are Prometheus's shape — a Grafana over them is the dashboard.

NACRE_URL        the API's base URL
NACRE_TOKEN      a key holding write on its layers
CONNECTOR_STATE  /state/connector.sqlite
SYNC_INTERVAL    300 seconds, at least 10
SYNC_ONCE        true: one sweep, exit with its verdict
PORT             9400

What they are not

Not a crawler, and not every source. There is no Confluence, SharePoint, Slack, Notion or Jira connector. Content there is pushed through the API by whoever owns the sync, which is the arrangement Nacre was built around in the first place.

Files — PDF, Word, PowerPoint, Excel, OpenDocument, EPUB, RTF — go up as files, and the index extracts the text, so the deployment needs S3-compatible object storage to accept them. Text needs nothing extra.

The images are on ghcr.io/nacre-work/connectors/<name>, for amd64 and arm64, and mirrored to Docker Hub as nacrecontextlayer/connector-<name>. Contributions are covered by the same contributor agreement as the core.