diff --git a/docs/agent-toolkit/mcp-server.md b/docs/agent-toolkit/mcp-server.md index 0be995bb..bdc97713 100644 --- a/docs/agent-toolkit/mcp-server.md +++ b/docs/agent-toolkit/mcp-server.md @@ -31,7 +31,7 @@ The MCP server does not receive your prompts and does not send any data to a thi Three things are enforced on the server before a result leaves epilot: - **Permissions.** Every tool runs with your epilot permissions. Upstream APIs apply their normal permission checks on every call, so the assistant can only see and change what you can. -- **PII anonymization.** Entity data on OAuth connections is anonymized server-side by an access token minted with the `anonymize` flag. The assistant cannot disable it. See [PII anonymization](/docs/agent-toolkit/setup#pii-anonymization). +- **PII anonymization.** Entity data on OAuth connections is anonymized server-side by an access token minted with the `anonymize` flag. The assistant cannot disable it. Detection is best effort; see [PII Anonymization](/docs/auth/anonymization) for what is covered. - **Credential redaction.** `auth` blocks, signed journey tokens, and portal authentication infrastructure are stripped from responses and reported in `redacted_fields`. ## Permissions diff --git a/docs/agent-toolkit/setup.md b/docs/agent-toolkit/setup.md index 87783ed1..ccbb85f7 100644 --- a/docs/agent-toolkit/setup.md +++ b/docs/agent-toolkit/setup.md @@ -297,7 +297,7 @@ Disconnecting also revokes the integration token that the connection created in ## PII anonymization -Entity data that the epilot MCP server returns to your AI assistant is **anonymized by default**. The OAuth connection mints an epilot access token with the `anonymize` flag set, which forces PII anonymization on all entity data returned to that token. The assistant cannot disable it. `whoami` reports this as `entity_pii: anonymized`. +Entity data that the epilot MCP server returns to your AI assistant is **anonymized by default**. The OAuth connection mints an epilot access token with the `anonymize` flag set, which forces PII anonymization on all entity data returned to that token. The assistant cannot disable it. `whoami` reports this as `entity_pii: anonymized`. Anonymization is [best effort](/docs/auth/anonymization): custom attributes are only masked when they are recognized as personal data or marked as PII in the schema. Anonymization is implemented by the open-source [@epilot/anonymization](https://github.com/epilot-dev/anonymization) library, the shared implementation behind epilot's anonymized API responses (`?anonymize=true` on the Entity API, or access tokens created with `anonymize: true`). It replaces personal data with deterministic pseudonyms, so the same person gets the same placeholder within an organization and the assistant can still reason about relations without seeing real values: @@ -331,12 +331,18 @@ If you need certainty for an attribute, classify it explicitly in the entity sch } ``` -Update the attribute in the [entity schema](/docs/entities/attributes) of your organization, for example through the Entity Builder or the Entity API, and the change applies to every anonymized response from then on, including everything the AI assistant reads. +Update the attribute in the [entity schema](/docs/entities/attributes#data-classification) of your organization, for example with the **Anonymize** checkbox in the Entity Builder or through the Entity API, and the change applies to every anonymized response from then on, including everything the AI assistant reads. :::tip Review your custom attributes once Before connecting an AI assistant, walk through the custom attributes of your contact, account, and order schemas and set `data_classification: "pii"` on every field that can contain personal data. The built-in defaults cover standard fields; your custom fields are where personal data slips through. ::: +:::caution Prefer read-only connections +Because the assistant only sees pseudonyms, a change it makes from what it has read can write pseudonyms back over your real data. The MCP server refuses generic API writes that contain masked values, but connect read-only (`?access=read`, or keep the preselected **Read-only** on the approval screen) unless the task needs write access. +::: + +For the full picture, including which APIs anonymize, what is not covered, and the classification rules, see [PII Anonymization](/docs/auth/anonymization). + ## Monitor usage - **Who is connected.** Every OAuth connection appears in epilot 360 under **Settings → Access Tokens** as an integration token named after the AI client and the approving user, for example `MCP: Claude (Erika)`. Deleting the token disconnects the assistant immediately. diff --git a/docs/auth/access-tokens.md b/docs/auth/access-tokens.md index 2de6e236..229b842e 100644 --- a/docs/auth/access-tokens.md +++ b/docs/auth/access-tokens.md @@ -35,6 +35,24 @@ When creating a token, you can optionally set an **expiry**. A token with an exp The generated token is shown only once and must be saved by the user. ::: +## Restricting Access Tokens + +Two options restrict what a token can do, independently of its roles. Both are set when the token is created, are enforced server-side, and can't be switched off by the token holder. + +### Read-only tokens {#read-only-tokens} + +A token created with **Read-only** (`read_only: true`) can perform read actions only. Every write action is denied, regardless of the token's roles. + +### Anonymized tokens {#anonymized-tokens} + +A token created with **Anonymize** (`anonymize: true`) receives PII-anonymized entity data: names, emails, phone numbers, addresses, IBANs, and other personal data are replaced with pseudonyms in every Entity API response. Entity exports are blocked for anonymized tokens. + +:::caution Anonymization is best effort +Standard fields are masked automatically, but custom attributes are only masked when they are recognized as personal data or marked **Anonymize** in the entity schema. Read [PII Anonymization](/docs/auth/anonymization) to learn what is covered and how to classify your attributes before you hand an anonymized token to an AI assistant or a third party. +::: + +Combine both options for AI assistants and analytics integrations: an anonymized token can't write pseudonyms back over your real data if it is also read-only. + ## Revoking Access Tokens Delete an Access Token from the management view to revoke it. After revocation, the token is immediately invalidated. @@ -82,6 +100,19 @@ Set an optional expiry with the `expires_in` parameter — a number of seconds ( } ``` +Set `read_only: true` or `anonymize: true` to create a [restricted token](#restricting-access-tokens): + +```json title="Request body for a read-only, anonymized token" +{ + "name": "AI analytics", + "read_only": true, + "anonymize": true, + "expires_in": "7d" +} +``` + +A token created by an anonymized token is always anonymized too, whatever the request body says. + Tokens created with `expires_in` are stored, listed, and revocable exactly like non-expiring tokens. The response includes an `expires_at` timestamp, and the token stops working — and drops out of the token list — once it expires: ```json title="201 response for a token with expiry" @@ -123,5 +154,6 @@ DELETE /v1/access-tokens/api_5ZugdRXasLfWBypHi93Fk ## See Also - [Token Types](/docs/auth/token-types) — comparison of all epilot token types +- [PII Anonymization](/docs/auth/anonymization) — how anonymized tokens mask personal data - [Authentication](/docs/auth/authentication) — OAuth 2.0 login flow - [Permissions](/docs/auth/permissions) — role-based access control and grants diff --git a/docs/auth/anonymization.md b/docs/auth/anonymization.md new file mode 100644 index 00000000..a6815e2e --- /dev/null +++ b/docs/auth/anonymization.md @@ -0,0 +1,163 @@ +--- +sidebar_position: 2.5 +title: PII Anonymization +--- + +# PII Anonymization + +[[Library source](https://github.com/epilot-dev/anonymization)] +[[npm](https://www.npmjs.com/package/@epilot/anonymization)] + +epilot can mask personal data (PII) in API responses, so that AI assistants, analytics, and integrations get the shape and relations of your data without the personal identifiers in it. This page explains how anonymization works, what it covers, where its limits are, and how you control it through your entity schemas. + +:::caution Anonymization is best effort +Anonymization reliably masks standard fields such as names, emails, phone numbers, addresses, and IBANs. It **cannot know** which of your custom attributes contain personal data. A custom attribute with an unusual name, or a name typed into a generic text field, is returned **as is** unless you [mark the attribute as PII](#control-anonymization-in-the-entity-schema) in the entity schema. + +Treat anonymization as data minimization, not as a guarantee that no personal data leaves epilot. Before you connect an AI assistant or a third-party tool to an anonymized token, [review your custom attributes](#review-your-schemas). +::: + +## When anonymization applies + +Anonymization is turned on per request or per access token: + +| How | Who can turn it off | Used by | +| --------------------------------------------------------------------------------------- | --------------------------------------- | -------------------------------------------------------------------------------- | +| Access token created with `anonymize: true` | Nobody. The token can never see raw PII | [MCP server](/docs/agent-toolkit/setup#pii-anonymization) connections (always on), CLI sessions in **Anonymize mode**, your own [access tokens](/docs/auth/access-tokens#anonymized-tokens) | +| `?anonymize=true` query parameter (or `"anonymize": true` in an Entity API search body) | The caller, by leaving it out | Callers that want anonymized output from a normal token | + +The token flag is enforced server-side and is one-way: an anonymized token can only create further anonymized tokens, and no request parameter turns anonymization off. + +### Which APIs anonymize + +| API | What is anonymized | +| ------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | +| [Entity API](/api/entity) | All entity data in successful JSON responses: get, search, relations, activity feed, and autocomplete. Uses the entity schema to classify every attribute. Anonymized responses carry the header `x-epilot-anonymized: true`. | +| [Audit Log API](/docs/audit-logs) | Audit log entries. There is no schema for log payloads, so detection relies on field names and on email, phone, and IBAN patterns. | +| APIs that read entities with your token | Inherit anonymization from the Entity API, for example email templates that resolve entity variables. | + +Entity exports (`exportEntities`) are **blocked** for anonymized callers, because export files are written outside the anonymized response pipeline. + +:::warning Other APIs are not anonymized +APIs that keep their own data, rather than reading it from the Entity API with your token, return their data unchanged, even to an anonymized token. Combine anonymization with a [read-only](/docs/auth/access-tokens) token and least-privilege roles, so a token can only reach the data it needs. +::: + +## How values are masked + +Anonymization is implemented by the open-source library [`@epilot/anonymization`](https://github.com/epilot-dev/anonymization). The complete list of what is classified as PII lives in its [`defaults.ts`](https://github.com/epilot-dev/anonymization/blob/main/src/defaults.ts) as plain, reviewable data. + +Values are replaced with **deterministic pseudonyms**: the same person gets the same pseudonym across requests within your organization, so an assistant can still group, join, and deduplicate records without seeing the real values. Pseudonyms differ between organizations and cannot be reversed without epilot's secret. + +| Data | Anonymized value | +| --------------------------------------------------------------- | ----------------------------------------------------------------------- | +| Names (`first_name`, `last_name`, contact title, …) | `person_4f2a9b1c` | +| Company names, tax and registration IDs | `company_4f2a9b1c`, `value_4f2a9b1c` | +| Email, phone, IBAN (by attribute type or field name) | Format-preserving: `4f2a9b1c@anonymized.invalid`, `+0083920174`, `iban_7c01d2aa` | +| Addresses | Postal code, city, and country kept; street pseudonymized; the rest dropped | +| Birthdates | Truncated to the year: `1985-06-14` → `1985-01-01` | +| Payment methods | IBAN and account holder pseudonymized; bank name and BIC kept | +| User relations (for example `contact_owner`) | Name and email pseudonymized, credentials removed, IDs kept | +| Consents | Email or phone pseudonymized; consent status and history kept | +| Known free text (note and comment content, message subject and body, opportunity description) | `[REDACTED]` | +| Emails, phone numbers, and IBANs **inside** other string values | Replaced with pseudonyms | +| Everything else (IDs, status, dates, numbers, unclassified text) | **Unchanged** | + +### How an attribute is classified + +For each attribute, the first matching rule wins: + +1. `data_classification: "public"` on the schema attribute: never anonymized. +2. `data_classification: "pii"` on the schema attribute: always anonymized. +3. A curated field name, such as `first_name`, `birthdate`, or `iban` on any schema, or `note:content`. +4. The attribute type: `email`, `phone`, `address`, `payment`, `relation_user`, and `consent` attributes. +5. Otherwise the value is kept, with emails, phone numbers, and IBANs embedded in it replaced. + +Rules 3 to 5 are the **best-effort** part. They cover epilot's standard attributes, but they don't know what your custom attributes mean. + +## What is not anonymized + +Personal data can still come through in these cases: + +- **Custom attributes** that don't match a curated name or a PII type, for example a `string` attribute `ansprechpartner` containing a name. Only emails, phone numbers, and IBANs typed into such a field are caught. +- **Free text** in custom text fields, such as a name or a customer ID written into an internal remark. +- **Values in unusual formats**, for example a phone number without a country prefix. +- **Other APIs** that do not read their data through the Entity API. See [Which APIs anonymize](#which-apis-anonymize). +- **Configuration data**, such as journey texts, workflow names, email template bodies, or automation settings. This isn't entity data and is returned as configured. +- **Combinations of kept values.** Postal code, city, birth year, and consumption data together can narrow down a person in small populations. +- **Search by value.** A caller who already knows an email address can still search for it and confirm that a matching record exists. + +Anonymization is pseudonymization in the sense of GDPR Art. 4(5): anonymized data is still personal data from a legal perspective. + +## Control anonymization in the entity schema + +Every entity schema attribute supports a `data_classification` property. It takes precedence over all built-in defaults. + +| `data_classification` | Effect | +| --------------------- | ----------------------------------------------------------------------------------------------------------------------------------- | +| `pii` | Always anonymized. Use it for every custom attribute and free-text field that can contain personal data. | +| `public` | Never anonymized, not even the embedded email, phone, and IBAN scrubbing. Use it for fields the defaults mask by mistake, such as a non-personal ID. | +| unset | Built-in best-effort defaults apply. | + +For an attribute marked `pii`, the value is masked based on its type: + +- Email, phone, address, and payment attributes use their type-specific masking. +- Dates are truncated to the year. +- Multiline text is replaced with `[REDACTED]`. +- Other single-line values get a stable pseudonym, such as `value_4f2a9b1c`, so they stay joinable. + +### In the Entity Builder + +Open the attribute in the Entity Builder and check **Anonymize** (German: **Anonymisieren**, "Mask this value in anonymized API responses") in its configuration. This sets `data_classification: "pii"`. Unchecking it removes the classification, and the built-in defaults apply again. To set `public`, use the Entity API. + +The checkbox is available for all attribute types except sequences and headlines. + +### Through the Entity API + +Set `data_classification` on the attribute when you update the schema, for example with [`putSchema`](/api/entity#tag/schemas/operation/putSchema): + +```json title="Marking a custom attribute as PII" +{ + "name": "internal_remarks", + "label": "Internal remarks", + "type": "string", + "multiline": true, + "data_classification": "pii" +} +``` + +```json title="Opting a non-personal identifier out" +{ + "name": "customer_number", + "label": "Customer number", + "type": "string", + "data_classification": "public" +} +``` + +The change applies to every anonymized response from then on, including everything an AI assistant reads through the MCP server. You don't have to recreate tokens. + +### Review your schemas + +Before you connect an AI assistant, the CLI in **Anonymize mode**, or an integration to an anonymized token: + +1. List the custom attributes of each schema that holds personal data, typically contact, account, opportunity, order, contract, meter, and your custom entities. +2. Mark every attribute that can contain personal data with **Anonymize**, including free-text fields such as remarks or descriptions. +3. Check the result: fetch an entity with `?anonymize=true` and look for real values. + +```bash +epilot entity getEntity contact -p anonymize=true +``` + +## Don't edit data through an anonymized connection + +An anonymized token sees pseudonyms instead of real values. If a tool writes an entity it has read back, it stores the pseudonyms in place of the real values, for example `4f2a9b1c@anonymized.invalid` instead of the customer's email address. + +- Use anonymized connections for reading and analysis. Create them **read-only** where possible. +- The epilot MCP server refuses writes through its generic API route that contain pseudonyms or `[REDACTED]` placeholders, but other clients don't check this. +- To change data, update only the fields you intend to change, with values you know are real. + +## Related + +- [Access Tokens](/docs/auth/access-tokens#anonymized-tokens): create anonymized tokens +- [CLI authentication](/docs/cli/overview#anonymized-sessions): anonymized CLI sessions +- [MCP server setup](/docs/agent-toolkit/setup#pii-anonymization): anonymization for AI assistants +- [Entity attributes](/docs/entities/attributes#data-classification): the `data_classification` property diff --git a/docs/cli/overview.md b/docs/cli/overview.md index c18c46f4..c3550ae7 100644 --- a/docs/cli/overview.md +++ b/docs/cli/overview.md @@ -101,6 +101,9 @@ epilot auth login # Browser-based login restricted to a read-only session epilot auth login --readonly +# Browser-based login with PII-anonymized entity data +epilot auth login --readonly --anonymize + # Manual token epilot auth login --token @@ -146,6 +149,20 @@ vs. a normal read-write token: Access: read-write ``` +### Anonymized sessions + +Pass `--anonymize` to `epilot auth login` to obtain an [anonymized token](/docs/auth/access-tokens#anonymized-tokens). Personal data in entity responses, such as names, emails, phone numbers, and addresses, is replaced with pseudonyms, and entity exports are blocked. Use it when the output of the CLI goes to an AI assistant or into logs. + +```bash +epilot auth login --readonly --anonymize +``` + +As with `--readonly`, the browser authorize page pre-checks and locks the **Anonymize mode** option. You can also check it manually during a normal `epilot auth login`. `epilot auth status` shows `Data: anonymized` for such a session. + +:::caution Anonymization is best effort +Custom attributes are only masked when they are recognized as personal data or marked **Anonymize** in the entity schema, and APIs other than the Entity API and Audit Log API return their data unchanged. See [PII Anonymization](/docs/auth/anonymization) for what is covered. Combine `--anonymize` with `--readonly`, so that pseudonyms can't be written back over real data. +::: + ## Parameters ```bash diff --git a/docs/entities/attributes.md b/docs/entities/attributes.md index bb76b84e..2c5cfd97 100644 --- a/docs/entities/attributes.md +++ b/docs/entities/attributes.md @@ -34,6 +34,29 @@ All attribute types share these base properties: | `repeatable` | `boolean` | Allow multiple values (see [Repeatable Attributes](#repeatable-attributes)) | | `has_primary` | `boolean` | Support marking one item as primary | | `render_condition` | `string` | Conditional visibility expression (see [Conditional Rendering](#conditional-rendering)) | +| `data_classification` | `"pii"` \| `"public"` | Controls PII anonymization of the value (see [Data Classification](#data-classification)) | + +### Data Classification + +`data_classification` tells [PII anonymization](/docs/auth/anonymization) whether an attribute contains personal data. It applies to anonymized responses, meaning requests with `?anonymize=true` and tokens created with `anonymize: true`, such as MCP server connections. + +| Value | Effect | +|-------|--------| +| `pii` | The value is always anonymized. Set it on custom attributes and free-text fields that can contain personal data. | +| `public` | The value is never anonymized. Set it on fields the built-in defaults mask by mistake. | +| unset | Best-effort defaults based on the attribute type and a curated list of field names. Custom attributes are usually **not** masked. | + +```json +{ + "name": "internal_remarks", + "label": "Internal remarks", + "type": "string", + "multiline": true, + "data_classification": "pii" +} +``` + +In the Entity Builder, check **Anonymize** on the attribute to set `pii`. ---