Skip to main content
Background sync collects content from connected services, such as their Notion pages or Google Drive files, and returns it as one uniform JSON export. You save a sync config that names the servers, the users, and what to collect. Then you trigger it from your backend whenever you want fresh content for a search index, RAG pipeline, or knowledge base. Each trigger starts one background agent per user and server. The agent finds and reads the content your instructions describe, as that user, and decides which resources belong in the export. The export contains the content exactly as the service returned it, not an agent’s summary of it.
Background sync is a premium feature. A tenant admin turns it on under Administration > Premium Features. Creating, editing, and triggering configs requires it. Reading configs, runs, transcripts, and exports, and deleting configs, keep working when it’s off.

Concepts

How It Works

You call every Background Sync endpoint with your platform access token; you don’t need a navigator or a navigator API key. Sync agents are instructed to only read, and tools that require approval are never available to them.

Supported servers

Background Sync works with selected servers from the Caylex catalog. These are supported today:
NotionGoogle DriveGoogle DocsGitHub
Call GET /sync-configs/sources to see which servers in your project can be synced. Custom servers aren’t supported.

How the agent decides what to export

While it crawls, the agent records a decision for every resource it reads. You’ll see these decisions in the run’s transcript:
  • Include a resource that matches your instructions. Including a container, such as a folder, parent page, or database, with its descendants covers everything under it that the run fetched.
  • Exclude a resource that doesn’t match, with a reason. Excluding a container also excludes what’s only reachable through it.
  • Unresolved marks a request it couldn’t fulfil, for example because the service needs authentication.
Only included resources are exported. Anything the agent read but didn’t decide on is listed in the export’s undecided array instead.

Before You Start

You need:
  • the Background Sync toggle turned on under Administration > Premium Features in the Caylex Platform (only a tenant admin can change it);
  • a server-side platform access token;
  • a project with at least one supported, unpaused server; and
  • users of that project who have authenticated with those servers. Project-level shared credentials also count.
All endpoints use the Platform API base URL and Bearer authentication:
Exports and transcripts contain complete customer documents, database rows, comments, and other sensitive content. Keep the platform token and everything you download server-side, apply your normal retention controls, and never commit exports to source control.

1. Find the Servers You Can Sync

This returns the project’s unpaused servers that support Background Sync:

2. Create a Sync Config

The response is the saved config, with each server’s source_type and an available flag. available turns false if a server is later paused or stops supporting Background Sync, and triggers skip that server. Creating a config returns 400 for an invalid server or user list, 404 when the project doesn’t exist, and 409 when the project already has a config with that name.

Writing good instructions

The agent searches and reads based on your instructions, so be as specific as you would with a colleague:
  • Name the scope: the workspace, teamspace, folder, repository, or page tree to start from.
  • Say what counts: document types, topics, owners, or labels that belong in the export.
  • Say what doesn’t: archived, draft, or personal content you want left out.
Use the trigger’s since field for date ranges instead of writing them into the instructions, so the same config works for full and incremental syncs.

3. Check User Access

Before you trigger, you can see which users can reach which servers:
A trigger skips every user and server pair that isn’t accessible: project_user: false flags a listed user who has left the project.
Google Drive and Google Docs sources can read only the files each user has selected to share with Caylex. Content outside that selection won’t appear in their exports.

4. Trigger a Sync

The request accepts only these fields. Every run uses the config’s saved instructions. The trigger starts every run, or none of them:
Each source’s source label is unique within its run. It’s the source type, qualified by the server name when a config includes two servers of the same type, for example notion-team-wiki. Use it to address one source in the status, export, transcript, and continue endpoints. config_warnings lists changes to the server’s tools that may affect what the sync can collect; the source still runs. skipped lists the first 100 user and server pairs that didn’t run, and skipped_count has the total: A trigger returns 400 when no user can access any of the config’s servers, or when it would exceed the limits. It returns 409 when Caylex can’t start the sync; detail says why.

5. Poll for Status

The response above is trimmed. Each source also repeats its source_type, server_name, server_instance_id, resolved_by, and config_warnings, and summary carries the full counts and diagnostics described in Export format. summary is set once the source finishes. A run’s status summarizes its sources: Each source’s status is PENDING, RUNNING, COMPLETED, INCOMPLETE, or FAILED. A source is INCOMPLETE when the agent stopped before it finished, for example because it ran out of iterations, or when it read content but included none of it and left some undecided. A FAILED source hit an error; see error.

Track the runs in a trigger

A trigger has no status of its own. Its trigger_id only groups the runs it created, one per user, and each run finishes independently. The status in the trigger response is always pending and isn’t updated afterwards. Use a run’s sync_id to check or act on that user’s sync. The status, export, transcript, and continue endpoints all take a sync_id. Passing a trigger_id to them returns 404. Keep each sync_id from the trigger response’s runs, along with its user_email. To check every run in a trigger with one call, filter the run list by trigger_id:
Each item is one run, with the run and source status values described above. The response is trimmed: each item also has the run’s instructions, since, and created_at. The list doesn’t include a source’s summary, error, or timestamps; call GET /sync-task/{sync_id} for those. Runs are listed newest first. A trigger can create up to 500 runs and size is at most 100, so pass meta.next_cursor as cursor while meta.has_next is true. The list also gives you each run’s sync_id again if you didn’t keep the trigger response. You don’t need to wait for the whole trigger. As each run reaches completed, incomplete, or failed, use its sync_id to fetch its export, continue it if it’s incomplete, or read its sources’ error with GET /sync-task/{sync_id} if it failed. The trigger is done when every run has finished. This loop does that with the client from the example:
Here handle_run is your own code that fetches the export, continues the run, or records the failure. You can also filter GET /sync-task by project_id or sync_config_id to list runs across triggers.

6. Fetch the Export

Without source, the export covers every source and is available once the run’s status is completed, incomplete, or failed. Until then it returns 409, with each source’s status in detail.sources. Resources are sorted by key. Pass page.next_offset as offset until it’s null. The excluded, undecided, unresolved_requests, and unmatched_tool_calls arrays are returned only on the first page (offset=0).
An export is available only while Caylex retains the run’s data. Pull it once the run finishes and store it in your own system rather than treating Caylex as its store.

Export Format

Resources

Sources

Each entry in sources summarizes one source. Its status: counts breaks down what happened: The remaining top-level arrays explain what isn’t in resources. excluded lists the agent’s exclusions with reasons. undecided lists content it read but never decided on. unresolved_requests lists what it couldn’t collect, for example {"request": "...", "reason": "authentication_required"}. unmatched_tool_calls lists tool calls Caylex couldn’t attribute to a resource, each with a reason. schema_version changes only when the export shape changes in a breaking way. engine_version changes when Caylex changes how it builds resources, so two exports of the same run can differ across versions.

Incremental Syncs

Pass since when you trigger to collect only what changed:
  1. Record the time you trigger.
  2. On the next trigger, pass the previous trigger’s time as since.
  3. Upsert each exported resource by key, skipping those whose content_hash matches your stored copy.
With since set, the export drops included resources modified before it and resources that have no modified time, and counts them in skipped_unchanged. The export’s top-level since is the cutoff that was applied. It’s null only when no source was date-filtered, so you can treat a null export as a full snapshot. An export lists what the run collected; it doesn’t report deletions. To find resources that were deleted or that a user lost access to, run a periodic full sync without since and compare its keys with your store.
Each run collects as one user, with that user’s access. The same resource can appear in several users’ exports under the same key. Deduplicate by key, but keep track of which users’ runs returned it if your index enforces per-user access.

Continue a Run

When a source stopped short, or you want more from one that finished, send the agent a follow-up message in the same session:
Continuing resumes the source with its transcript intact, so the agent keeps what it already collected and adds to it. Poll the run again and fetch a fresh export when it finishes. Runs stay continuable after you delete their config. The endpoint returns 409 when there’s nothing to continue or the named source is still running or failed.

Read a Transcript

To see what the agent did, read one source’s transcript, oldest message first:
source is required when the run has more than one source. The response has the source’s task_id and status, its messages (each with sequence_number, role, message_type, content, and created_at), and next_cursor and has_more. Pass next_cursor back as cursor to read the next page while has_more is true. limit defaults to 50 messages, up to 100.

Manage Sync Configs

Editing a config affects only future triggers. Runs that already started keep the settings they were triggered with.

Limits

Example

This example triggers a config, waits for every run to finish, and pages through each run’s export.
background_sync.py