Commit Graph
8 Commits
Author SHA1 Message Date
Theodore LiandClaude Fable 5 bcad4e308c fix(pii): mask offloaded large payloads chunk-by-chunk instead of aborting at 16MB (#5810)
* fix(pii): mask offloaded large payloads chunk-by-chunk and retry transient mask failures

A block output past the 16MB inline materialization ceiling aborted the run
before masking even started: the redaction path hydrated the whole offloaded
value at once, and the pre-flight size assert fired on the manifest's total
byteSize. Large-array manifests now page one stored chunk at a time
(materialize -> mask -> re-store, rebuilt via the manifest writer with preview
derived from masked items), so peak heap stays ~one chunk regardless of payload
size. Single refs up to the 64MB durable cap hydrate with a raised budget and
run serially outside the concurrency pool.

Mask-batch chunk requests now retry transient failures (network errors,
408/429/5xx, honoring Retry-After) with jittered backoff, so a single ALB blip
or Presidio pod restart no longer fails a whole payload's redaction. Nested-ref
masking now runs the string pass before ref substitution, fixing a latent
double-mask when a masked nested value shrinks back inline.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A1JYstmLHk9qMGyBDqYRcJ

* fix(pii): retry runtime timeouts and socket closes in mask-batch chunks

Verified end-to-end against a 26MB / 40k-record offloaded output: the
chunk-wise path masks it in ~54s on a single local Presidio worker where the
old path aborted at the 16MB ceiling. The exercise surfaced two more transient
error shapes the retry classifier missed — runtime-level request timeouts
(undici's default 300s headers timeout, Bun's TimeoutError) and mid-flight
socket closes — both of which previously failed the whole payload's redaction
on the first occurrence.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A1JYstmLHk9qMGyBDqYRcJ

* fix(pii): gate the spaCy fast path on entities the loaded models can produce

The registry's SpacyRecognizers claim every entity in Presidio's default
NER-model mapping — including PHONE_NUMBER/AGE/ID/EMAIL, which exist for
transformer de-identification backends and which no spaCy model can emit. The
NER_ENTITIES derivation trusted that claim, so any request naming PHONE_NUMBER
(present in nearly every redaction rule) silently forced the full spaCy pass
and the regex-only fast path never fired. Intersect the claimed set with the
entities the loaded models' actual NER labels map onto; the hard floor of core
NER entities is unchanged, and a future backend that genuinely emits phone
labels would re-gate automatically.

Verified live: PHONE_NUMBER-only requests take nlp=skip with span parity
against the full path, PERSON still forces NER, and a 26MB/40k-record
block-output redaction with the realistic entity set runs entirely on the
fast path (~3.2min vs ~15min projected full-NER on one worker).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A1JYstmLHk9qMGyBDqYRcJ

* fix(pii): hydrate oversized-chunk manifests serially, not just single refs

Review finding: a manifest whose packer emitted a chunk past the inline
ceiling (one item larger than the chunk target) hydrates that chunk with the
raised 64MB budget inside the REF_CONCURRENCY pool, so several such manifests
could hydrate oversized blobs concurrently — the exact heap scenario the
serial path exists to prevent. The serial gate now covers any ref whose
hydration can exceed the inline ceiling: oversized single refs and manifests
containing an oversized chunk. Normally-chunked manifests stay pooled.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A1JYstmLHk9qMGyBDqYRcJ

* fix(pii): serialize oversized hydrations globally across nested redaction passes

Review finding: the serial gate was per-resolveReplacements invocation, so
nested oversized refs discovered inside different pooled parents each got
their own pool and could hydrate oversized blobs concurrently. A shared
promise-chain gate now threads through the options from the entry points, and
a reentrancy flag lets a gated ref's own nested oversized work run directly
instead of deadlocking on the hold. Covered by a cross-parent max-in-flight
assertion and a nested-oversized deadlock regression test.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A1JYstmLHk9qMGyBDqYRcJ

* fix(pii): fail fast on a null mask-batch body; use sleep() in gate test

A 200 response with a null JSON body threw TypeError on the data.masked read,
which the retry classifier treats as transient — burning the full retry budget
on a deterministic shape failure. Null-guard the body so it throws the
non-retryable shape error immediately. Also swap the gate test's inline
setTimeout promise for sleep() to satisfy check:utils, which failed CI.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A1JYstmLHk9qMGyBDqYRcJ

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 18:03:14 -04:00
Theodore Li 61de0b5808 feat(pii): custom user-supplied regex patterns for redaction (#5732)
* feat(pii): custom user-supplied regex patterns for redaction

* fix(pii): enforce custom-regex syntax + safety at the boundary schema

* improvement(pii): always wrap custom-pattern redaction token in angle brackets

* chore(pii): register guardrails_validate in the dev minimal tool registry

* fix(pii): coerce empty guardrails entity-type checkbox (null) so the contract accepts it

* fix(pii): keep detect-all when a custom pattern is added; custom patterns win overlaps
2026-07-17 15:20:58 -04:00
Theodore Li 4d6301c900 feat(pii): regex-only block-output redaction + drop GLiNER/GPU image (#5697)
* chore(pii): remove GLiNER/GPU image + add spaCy-skip fast path to CPU server

* feat(pii): restrict block-output redaction to regex-only entities

* fix(pii): derive spaCy-NER set from registry + skip fast path when score_threshold set

* fix(pii): include ORGANIZATION in app-side NER set (align with server)
2026-07-15 20:28:15 -04:00
Theodore LiandClaude Fable 5 4e6594dc54 feat(pii): add opt-in GLiNER NER engine (PII_ENGINE), device-agnostic (#5495)
* feat(pii): add opt-in GLiNER NER engine (PII_ENGINE), device-agnostic

Swap the 4 NER entity types (PERSON/LOCATION/NRP/DATE_TIME) to a single
multilingual GLiNER zero-shot model when PII_ENGINE=gliner; spaCy stays the
default and all ~36 regex/checksum recognizers are identical on both engines.
Device-agnostic via PII_DEVICE / cuda auto-detect — same code on Fargate CPU
now and EC2-GPU later.

- engines.py: side-effect-free builders; SharedModelGLiNERRecognizer loads
  ONE model shared across the 5 per-language instances and restricts labels
  to the entities it owns; small spaCy models keep tokenization/lemmas for
  the regex recognizers; fail-fast on the lean image
- pii.Dockerfile: multi-stage — default target unchanged (lean spaCy);
  --target gliner is a superset (torch CPU + gliner + baked model) where
  both engines work; gliner-gpu scaffold for the GPU fleet
- CI publishes the gliner variant (:staging-gliner/:latest-gliner, amd64)
- Helm: pii.engine / pii.device values wired to PII_ENGINE/PII_DEVICE
- scripts/bench_engines.py: throughput + NER-parity diff harness
- tests: unit (mocked GLiNER) + in-image integration for both engines

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Up3F97mjCH9HCj1pX4J8VJ

* refactor(pii): ship both engines in one image — engine is a pure env flip

Collapse the gliner build target into the single pii image: spaCy lg models,
torch (CPU), gliner, and the baked GLiNER weights all ship in it, so
PII_ENGINE switches engines with no image swap and no tag matrix. CI reverts
to the single pii build (no -gliner tags). The GPU variant becomes the same
Dockerfile built with --build-arg TORCH_INDEX_URL=.../cu128. Image grows
~6.1GB -> ~9.6GB.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Up3F97mjCH9HCj1pX4J8VJ

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-07 22:37:33 -04:00
Theodore Li 34b2abadcb perf(pii): mask offloaded refs + parallelize + raise redaction concurrency (#5436)
* perf(pii): mask offloaded refs in block outputs/input + parallelize + raise concurrency

- Block-output and input stages now hydrate → mask → re-store large-value refs
  (function/tool outputs are offloaded to refs before reaching the redactor, which
  treats refs as opaque) — fixes big block outputs coming back unredacted
- Parallelize large-value ref hydrate/mask/re-store: collect refs across the whole
  payload and process them with bounded concurrency instead of one sequential pass
  per key (PII_REF_CONCURRENCY, default 4)
- Raise mask-batch chunk concurrency default 4 -> 64 to saturate the load-balanced
  Presidio fleet behind the internal ALB (PII_MASK_CHUNK_CONCURRENCY, env-tunable)

* fix(pii): keep resolveReplacements mapper total (respect mapWithConcurrency contract)

- Catch per-ref errors inside the mapWithConcurrency mapper and rethrow the first
  after the pool drains, instead of letting a throwing mapper reject the pool
  (the util documents fn MUST NOT reject). Preserves fail-fast abort in throw mode.
- Add multi-ref throw-mode test: one of several refs failing aborts the redaction

* feat(pii): make route->Presidio chunk concurrency env-tunable (PII_SERVICE_CHUNK_CONCURRENCY)

Was hardcoded to 4; now env-tunable (default 4) so the inner route->Presidio
fan-out can scale with the fleet alongside the outer PII_MASK_CHUNK_CONCURRENCY.

* feat(pii): combined /redact + /redact_batch endpoint (one round-trip)

- Presidio: add /redact and /redact_batch (analyze + anonymize server-side, feeding
  analyzer results straight to the anonymizer, no dict round-trip). All existing
  endpoints kept, so old clients keep working during rollout.
- App: maskPIIBatch now uses /redact_batch (halves app<->service round-trips, no
  shipping text back up for anonymize). Falls back to the legacy analyze_batch +
  anonymize_batch pair on a 404 (older Presidio image), so the app is safe to deploy
  before or after the service in either order. Same length-check fail-closed guarantee.
2026-07-06 16:33:44 -04:00
Theodore Li 69b81a679b feat(data-retention): granular PII redaction stages (input + block outputs) (#5272)
* feat(data-retention): granular PII redaction stages (input + block outputs)

* fix(data-retention): propagate block-output redaction into child workflows

* fix(data-retention): close block-output redaction gaps on streaming + resume

* fix(data-retention): drain+mask streamed output, resolve PII policy unconditionally (no fail-open)

* test(testing): support leftJoin().where().limit() in shared db mock

* fix(data-retention): mask agent/Pi memory writes under block-output redaction

* fix(data-retention): guard partial PII stages in GET normalize

* fix(data-retention): mask seeded memory messages under block-output redaction

* fix(guardrails): fail closed on misaligned Presidio batch responses

* fix(data-retention): enabled stage with no entity types redacts all (no fail-open)

* fix(data-retention): reject enabled stage with no entity types; empty = off everywhere

* docs(data-retention): note resume remask covers inline values only

* fix(data-retention): scrub offloaded large-value refs from logs when block-output redaction is off

* fix(data-retention): hydrate, mask, and re-store large-value refs in logs (preserve redacted content)

* fix(data-retention): always apply logs policy to large-value refs when logs stage is on

* perf(data-retention): drop redaction byte ceiling, parallelize chunks (env-tunable), remove request timeouts, sync large-value walk

* feat(data-retention): gate granular PII stages behind pii-granular-redaction flag

- New pii-granular-redaction feature flag (fallback PII_GRANULAR_REDACTION),
  layered on pii-redaction, gating the execution-altering input + block-output stages
- Route returns piiGranularRedactionEnabled and rejects enabling granular stages when off
- UI shows only the Logs stage tab unless the flag is on; clamps active stage
- Drop the per-search Select all toggle; add a Deselect all action to the PII section header

* docs(pii): describe Presidio as a standalone service, not a sidecar

Presidio now runs as its own ECS service (and, in Helm, its own Deployment +
Service) reached over the network via PII_URL — not a sidecar in the app task.
Update README, code comments, env docs, Dockerfiles, and the Helm chart docs to
match, and note the deploy requirement that PII_URL must be reachable.

* fix(data-retention): re-mask offloaded large-value refs on resume + don't lock out granular saves

- Resume/run-from-block restore now hydrates → masks → re-stores large-value refs
  in restored blockStates (not just inline strings), so a value offloaded before the
  block-output stage was enabled can't warm raw PII into downstream blocks. Fails fast.
- pii-large-values: add onFailure mode (throw on the execution path, scrub for logs)
  and redactLargeValueRefsInValue for arbitrary (non-RedactablePayload) values
- Granular flag gate now rejects only NEW off→on granular enablement, so orgs that
  already configured granular stages can still save retention settings when the flag is off
2026-07-01 21:47:02 -04:00
Theodore Li 23ec96b02a feat(pii): add redaction timing metrics across sidecar and persist path (#5264)
- Log per-request duration in the Presidio sidecar (/analyze, /anonymize)
- Add durationMs to the mask-batch endpoint log line
- Emit per-execution PII redaction timing (stringCount, totalBytes, durationMs, scrubbed)
2026-06-29 14:04:49 -04:00
Theodore Li 0191a614b6 feat(pii): build & own combined PII (analyzer + anonymizer) image (#5176)
* feat(presidio): build & own combined analyzer+anonymizer image

Replace the stock mcr.microsoft.com/presidio-* sidecar images with a single
image we build and push to ECR/GHCR. A thin FastAPI service constructs one
AnalyzerEngine + one AnonymizerEngine at startup and serves both on port 3000
(/health, /supportedentities, /analyze, /anonymize) so the app needs one
PRESIDIO_URL. English only; pinned presidio 2.2.362 + en_core_web_lg 3.8.0.

Bakes in the native check-digit VIN recognizer and registers 12 English
recognizers Presidio ships but does not load by default (UK_NINO, AU_*, IN_*,
SG_*), taking the supported English set from 19 to 32.

* feat(presidio): add multi-language support (es/it/pl/fi)

Configure a multi-language spaCy NLP engine (en/es/it/pl/fi lg models) and
explicitly register the national-id recognizers Presidio ships but does not
load by default: ES_NIF/NIE, IT_FISCAL_CODE/DRIVER_LICENSE/VAT_CODE/PASSPORT/
IDENTITY_CARD, PL_PESEL, FI_PERSONAL_IDENTITY_CODE. Verified the NLP-engine +
explicit-registration path detects in-language (Finnish id, score 1.0).

* improvement(presidio): address review feedback

- Register VIN under all served languages, not just en (Bugbot: VIN missed for
  non-English language routing).
- Bump HEALTHCHECK start-period to 180s — five lg models load at import (Bugbot).
- Drop --no-cache-dir so the pip cache mount actually works (Greptile).
- Pydantic request models for /analyze + /anonymize so missing 'text' returns
  422 not 500; default operator 'type' to 'replace' instead of KeyError->500
  (Greptile).

* refactor(pii): rename presidio image artifacts to pii

Rename the image/repo/secret/files from 'presidio' to 'pii' for clarity — the
service does PII detection + anonymization (and backs the guardrails block's
block/mask), not just redaction, and 'pii' matches existing pii-* naming.

docker/presidio.Dockerfile -> docker/pii.Dockerfile
docker/presidio/ -> docker/pii/
ghcr.io/simstudioai/presidio -> .../pii
ECR_PRESIDIO secret -> ECR_PII (infra side already renamed)
No behavior change — paths/identifiers only.

* refactor(pii): move service to apps/pii, make image ECR-only

- Move server.py + requirements.txt from docker/pii/ to apps/pii/ (source belongs
  under apps/, matching app/realtime; Dockerfile stays in docker/). Add a minimal
  @sim/pii package.json so the apps/* bun workspace glob accepts the Python service.
- Repoint docker/pii.Dockerfile COPY paths to apps/pii/; rename the container user
  presidio -> pii.
- Drop GHCR for pii — it's a private ECS sidecar pulled from ECR, never published.
  Removed it from the arm64/manifest (GHCR-only) jobs and guarded the build-amd64
  tag step to skip GHCR when no ghcr_image is set.
2026-06-23 03:41:35 -04:00