Commit Graph
9 Commits
Author SHA1 Message Date
Theodore Li 61de0b5808 feat(pii): custom user-supplied regex patterns for redaction (#5732)
* feat(pii): custom user-supplied regex patterns for redaction

* fix(pii): enforce custom-regex syntax + safety at the boundary schema

* improvement(pii): always wrap custom-pattern redaction token in angle brackets

* chore(pii): register guardrails_validate in the dev minimal tool registry

* fix(pii): coerce empty guardrails entity-type checkbox (null) so the contract accepts it

* fix(pii): keep detect-all when a custom pattern is added; custom patterns win overlaps
2026-07-17 15:20:58 -04:00
Theodore Li 4d6301c900 feat(pii): regex-only block-output redaction + drop GLiNER/GPU image (#5697)
* chore(pii): remove GLiNER/GPU image + add spaCy-skip fast path to CPU server

* feat(pii): restrict block-output redaction to regex-only entities

* fix(pii): derive spaCy-NER set from registry + skip fast path when score_threshold set

* fix(pii): include ORGANIZATION in app-side NER set (align with server)
2026-07-15 20:28:15 -04:00
Theodore LiandClaude Opus 4.8 97bb727eeb fix(pii): install CUDA torch on amd64 so GLiNER can run on GPU (#5552)
* fix(pii): install CUDA torch on amd64 so GLiNER can run on GPU

The published pii image installed a CPU-only torch build, so GLiNER on the
ECS GPU fleet died at model load with "Attempting to deserialize object on a
CUDA device but torch.cuda.is_available() is False". The Dockerfile already
had a TORCH_INDEX_URL arg, but no CI job ever passed --build-arg, so every
image silently took the cpu default.

Select the wheel index from TARGETARCH instead: amd64 gets cu128, arm64 keeps
the cpu index (cu128 publishes no aarch64 wheel at 2.11.0, and no arm64 target
has a GPU). CUDA torch falls back to CPU when no GPU is present, so one image
still serves both the Fargate CPU tasks and the EC2 GPU tasks off the same tag
— no CI or CDK changes needed.

cu128 keeps sm_75, the compute capability of the fleet's T4s, and its CUDA 12.8
runtime needs driver >=525 via minor-version compatibility, which the ECS GPU
AMI satisfies. cu121 was not an option: that index stops at torch 2.5.1.

Verified in an amd64 build of the changed block:
  2.11.0+cu128  cuda=12.8  arch=sm_75 sm_80 sm_86 sm_90 sm_100 sm_120
arm64 still resolves to 2.11.0+cpu. A build-time assert now fails the image
if amd64 ever silently regresses to a cpu wheel.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QHNEWVrh7k89m8Wtqzhs18

* fix(pii): assert torch CUDA state after every pip install

The check sat directly after the torch install, but requirements-gliner.txt
and requirements-dev.txt are installed afterwards and resolve against PyPI
with no torch pin, so a future gliner bump could swap the wheel that
torch_index selected without tripping the assert.

Neither file changes torch today (verified: torch is 2.11.0+cu128 both before
and after the gliner install), so this guards the invariant rather than fixing
a live regression. Moving it below the last pip install makes it certify the
torch that actually ships.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QHNEWVrh7k89m8Wtqzhs18

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-10 16:21:45 -04:00
Theodore LiandClaude Fable 5 4e6594dc54 feat(pii): add opt-in GLiNER NER engine (PII_ENGINE), device-agnostic (#5495)
* feat(pii): add opt-in GLiNER NER engine (PII_ENGINE), device-agnostic

Swap the 4 NER entity types (PERSON/LOCATION/NRP/DATE_TIME) to a single
multilingual GLiNER zero-shot model when PII_ENGINE=gliner; spaCy stays the
default and all ~36 regex/checksum recognizers are identical on both engines.
Device-agnostic via PII_DEVICE / cuda auto-detect — same code on Fargate CPU
now and EC2-GPU later.

- engines.py: side-effect-free builders; SharedModelGLiNERRecognizer loads
  ONE model shared across the 5 per-language instances and restricts labels
  to the entities it owns; small spaCy models keep tokenization/lemmas for
  the regex recognizers; fail-fast on the lean image
- pii.Dockerfile: multi-stage — default target unchanged (lean spaCy);
  --target gliner is a superset (torch CPU + gliner + baked model) where
  both engines work; gliner-gpu scaffold for the GPU fleet
- CI publishes the gliner variant (:staging-gliner/:latest-gliner, amd64)
- Helm: pii.engine / pii.device values wired to PII_ENGINE/PII_DEVICE
- scripts/bench_engines.py: throughput + NER-parity diff harness
- tests: unit (mocked GLiNER) + in-image integration for both engines

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Up3F97mjCH9HCj1pX4J8VJ

* refactor(pii): ship both engines in one image — engine is a pure env flip

Collapse the gliner build target into the single pii image: spaCy lg models,
torch (CPU), gliner, and the baked GLiNER weights all ship in it, so
PII_ENGINE switches engines with no image swap and no tag matrix. CI reverts
to the single pii build (no -gliner tags). The GPU variant becomes the same
Dockerfile built with --build-arg TORCH_INDEX_URL=.../cu128. Image grows
~6.1GB -> ~9.6GB.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Up3F97mjCH9HCj1pX4J8VJ

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-07 22:37:33 -04:00
Theodore Li 07a12249f5 feat(pii): env-driven uvicorn worker count (PII_WORKERS) (#5457)
* feat(pii): env-driven uvicorn worker count (PII_WORKERS)

Launch the Presidio service with --workers from the PII_WORKERS env var so one
image scales per task size (set PII_WORKERS = the task's vCPU count) without a
rebuild. `sh -c exec` expands the var while keeping uvicorn as PID 1 for clean
SIGTERM. Defaults to 1, so local/self-hosted is unchanged. Each worker loads the
spaCy models independently (~3.3GB measured), so task memory must be sized to
PII_WORKERS x ~3.3GB + overhead (set in infra alongside PII_WORKERS).

* harden(pii): quote PII_WORKERS expansion + raise healthcheck start-period for multi-worker

- Quote ${PII_WORKERS} so a malformed value fails uvicorn arg-parsing instead of
  being shell-interpreted (verified: '1; echo X' rejected as non-integer, X not run)
- Bump HEALTHCHECK start-period 180s -> 300s: N workers load the spaCy models in
  parallel, stretching cold start beyond the single-worker case
2026-07-06 21:23:48 -04:00
Theodore Li 69b81a679b feat(data-retention): granular PII redaction stages (input + block outputs) (#5272)
* feat(data-retention): granular PII redaction stages (input + block outputs)

* fix(data-retention): propagate block-output redaction into child workflows

* fix(data-retention): close block-output redaction gaps on streaming + resume

* fix(data-retention): drain+mask streamed output, resolve PII policy unconditionally (no fail-open)

* test(testing): support leftJoin().where().limit() in shared db mock

* fix(data-retention): mask agent/Pi memory writes under block-output redaction

* fix(data-retention): guard partial PII stages in GET normalize

* fix(data-retention): mask seeded memory messages under block-output redaction

* fix(guardrails): fail closed on misaligned Presidio batch responses

* fix(data-retention): enabled stage with no entity types redacts all (no fail-open)

* fix(data-retention): reject enabled stage with no entity types; empty = off everywhere

* docs(data-retention): note resume remask covers inline values only

* fix(data-retention): scrub offloaded large-value refs from logs when block-output redaction is off

* fix(data-retention): hydrate, mask, and re-store large-value refs in logs (preserve redacted content)

* fix(data-retention): always apply logs policy to large-value refs when logs stage is on

* perf(data-retention): drop redaction byte ceiling, parallelize chunks (env-tunable), remove request timeouts, sync large-value walk

* feat(data-retention): gate granular PII stages behind pii-granular-redaction flag

- New pii-granular-redaction feature flag (fallback PII_GRANULAR_REDACTION),
  layered on pii-redaction, gating the execution-altering input + block-output stages
- Route returns piiGranularRedactionEnabled and rejects enabling granular stages when off
- UI shows only the Logs stage tab unless the flag is on; clamps active stage
- Drop the per-search Select all toggle; add a Deselect all action to the PII section header

* docs(pii): describe Presidio as a standalone service, not a sidecar

Presidio now runs as its own ECS service (and, in Helm, its own Deployment +
Service) reached over the network via PII_URL — not a sidecar in the app task.
Update README, code comments, env docs, Dockerfiles, and the Helm chart docs to
match, and note the deploy requirement that PII_URL must be reachable.

* fix(data-retention): re-mask offloaded large-value refs on resume + don't lock out granular saves

- Resume/run-from-block restore now hydrates → masks → re-stores large-value refs
  in restored blockStates (not just inline strings), so a value offloaded before the
  block-output stage was enabled can't warm raw PII into downstream blocks. Fails fast.
- pii-large-values: add onFailure mode (throw on the execution path, scrub for logs)
  and redactLargeValueRefsInValue for arbitrary (non-RedactablePayload) values
- Granular flag gate now rejects only NEW off→on granular enablement, so orgs that
  already configured granular stages can still save retention settings when the flag is off
2026-07-01 21:47:02 -04:00
Theodore Li 76867062e5 feat(pii): publish PII image to GHCR and add Presidio sidecar to Helm chart (#5188)
* feat(pii): publish PII image to GHCR and add Presidio sidecar to Helm chart

* fix(pii): allow app→PII NetworkPolicy egress, global tolerations, topology spread
2026-06-23 21:07:06 -04:00
Theodore Li 4d2e7d5524 fix(pii): listen on 5001 to avoid app :3000 collision (awsvpc) (#5182)
* fix(pii): bind a configurable $PORT to avoid app :3000 collision

The pii image hardcoded uvicorn --port 3000 and ignored env. In the app ECS
task (awsvpc) all containers share one network namespace, and the app owns
3000 — so the sidecar must listen elsewhere (the stock presidio images honored
PORT and ran on 5002/5001). Bind ${PORT} (shell-form CMD), default 5001, and
update EXPOSE/HEALTHCHECK accordingly so the taskdef can set PORT=5001.

Verified: default binds 5001; PORT=5002 override binds 5002; /analyze works on
the overridden port.

* fix(pii): hardcode port 5001 (drop $PORT indirection)

EXPOSE can't be parameterized, so the configurable-PORT approach left EXPOSE
showing 5001 regardless (Greptile P2). We own both the image and the taskdef
and only ever need 5001, so hardcode it: exec-form CMD on 5001, EXPOSE 5001,
healthcheck on 5001. Runtime cmdline is identical to the verified ${PORT}
default (uvicorn ... --port 5001).
2026-06-23 04:46:21 -04:00
Theodore Li 0191a614b6 feat(pii): build & own combined PII (analyzer + anonymizer) image (#5176)
* feat(presidio): build & own combined analyzer+anonymizer image

Replace the stock mcr.microsoft.com/presidio-* sidecar images with a single
image we build and push to ECR/GHCR. A thin FastAPI service constructs one
AnalyzerEngine + one AnonymizerEngine at startup and serves both on port 3000
(/health, /supportedentities, /analyze, /anonymize) so the app needs one
PRESIDIO_URL. English only; pinned presidio 2.2.362 + en_core_web_lg 3.8.0.

Bakes in the native check-digit VIN recognizer and registers 12 English
recognizers Presidio ships but does not load by default (UK_NINO, AU_*, IN_*,
SG_*), taking the supported English set from 19 to 32.

* feat(presidio): add multi-language support (es/it/pl/fi)

Configure a multi-language spaCy NLP engine (en/es/it/pl/fi lg models) and
explicitly register the national-id recognizers Presidio ships but does not
load by default: ES_NIF/NIE, IT_FISCAL_CODE/DRIVER_LICENSE/VAT_CODE/PASSPORT/
IDENTITY_CARD, PL_PESEL, FI_PERSONAL_IDENTITY_CODE. Verified the NLP-engine +
explicit-registration path detects in-language (Finnish id, score 1.0).

* improvement(presidio): address review feedback

- Register VIN under all served languages, not just en (Bugbot: VIN missed for
  non-English language routing).
- Bump HEALTHCHECK start-period to 180s — five lg models load at import (Bugbot).
- Drop --no-cache-dir so the pip cache mount actually works (Greptile).
- Pydantic request models for /analyze + /anonymize so missing 'text' returns
  422 not 500; default operator 'type' to 'replace' instead of KeyError->500
  (Greptile).

* refactor(pii): rename presidio image artifacts to pii

Rename the image/repo/secret/files from 'presidio' to 'pii' for clarity — the
service does PII detection + anonymization (and backs the guardrails block's
block/mask), not just redaction, and 'pii' matches existing pii-* naming.

docker/presidio.Dockerfile -> docker/pii.Dockerfile
docker/presidio/ -> docker/pii/
ghcr.io/simstudioai/presidio -> .../pii
ECR_PRESIDIO secret -> ECR_PII (infra side already renamed)
No behavior change — paths/identifiers only.

* refactor(pii): move service to apps/pii, make image ECR-only

- Move server.py + requirements.txt from docker/pii/ to apps/pii/ (source belongs
  under apps/, matching app/realtime; Dockerfile stays in docker/). Add a minimal
  @sim/pii package.json so the apps/* bun workspace glob accepts the Python service.
- Repoint docker/pii.Dockerfile COPY paths to apps/pii/; rename the container user
  presidio -> pii.
- Drop GHCR for pii — it's a private ECS sidecar pulled from ECR, never published.
  Removed it from the arm64/manifest (GHCR-only) jobs and guarded the build-amd64
  tag step to skip GHCR when no ghcr_image is set.
2026-06-23 03:41:35 -04:00