mirror of
https://github.com/simstudioai/sim.git
synced 2026-09-24 15:45:35 +08:00
* feat(data-retention): granular PII redaction stages (input + block outputs) * fix(data-retention): propagate block-output redaction into child workflows * fix(data-retention): close block-output redaction gaps on streaming + resume * fix(data-retention): drain+mask streamed output, resolve PII policy unconditionally (no fail-open) * test(testing): support leftJoin().where().limit() in shared db mock * fix(data-retention): mask agent/Pi memory writes under block-output redaction * fix(data-retention): guard partial PII stages in GET normalize * fix(data-retention): mask seeded memory messages under block-output redaction * fix(guardrails): fail closed on misaligned Presidio batch responses * fix(data-retention): enabled stage with no entity types redacts all (no fail-open) * fix(data-retention): reject enabled stage with no entity types; empty = off everywhere * docs(data-retention): note resume remask covers inline values only * fix(data-retention): scrub offloaded large-value refs from logs when block-output redaction is off * fix(data-retention): hydrate, mask, and re-store large-value refs in logs (preserve redacted content) * fix(data-retention): always apply logs policy to large-value refs when logs stage is on * perf(data-retention): drop redaction byte ceiling, parallelize chunks (env-tunable), remove request timeouts, sync large-value walk * feat(data-retention): gate granular PII stages behind pii-granular-redaction flag - New pii-granular-redaction feature flag (fallback PII_GRANULAR_REDACTION), layered on pii-redaction, gating the execution-altering input + block-output stages - Route returns piiGranularRedactionEnabled and rejects enabling granular stages when off - UI shows only the Logs stage tab unless the flag is on; clamps active stage - Drop the per-search Select all toggle; add a Deselect all action to the PII section header * docs(pii): describe Presidio as a standalone service, not a sidecar Presidio now runs as its own ECS service (and, in Helm, its own Deployment + Service) reached over the network via PII_URL — not a sidecar in the app task. Update README, code comments, env docs, Dockerfiles, and the Helm chart docs to match, and note the deploy requirement that PII_URL must be reachable. * fix(data-retention): re-mask offloaded large-value refs on resume + don't lock out granular saves - Resume/run-from-block restore now hydrates → masks → re-stores large-value refs in restored blockStates (not just inline strings), so a value offloaded before the block-output stage was enabled can't warm raw PII into downstream blocks. Fails fast. - pii-large-values: add onFailure mode (throw on the execution path, scrub for logs) and redactLargeValueRefsInValue for arbitrary (non-RedactablePayload) values - Granular flag gate now rejects only NEW off→on granular enablement, so orgs that already configured granular stages can still save retention settings when the flag is off
51 lines
2.1 KiB
Docker
51 lines
2.1 KiB
Docker
# ========================================
|
|
# Combined Presidio service (analyzer + anonymizer) on a single port (5001)
|
|
# ========================================
|
|
FROM python:3.12-slim-bookworm AS base
|
|
|
|
WORKDIR /app
|
|
|
|
# build-essential for any sdist that compiles native deps (e.g. blis/thinc).
|
|
RUN --mount=type=cache,target=/var/cache/apt,sharing=locked \
|
|
--mount=type=cache,target=/var/lib/apt,sharing=locked \
|
|
apt-get update && apt-get install -y --no-install-recommends \
|
|
build-essential curl ca-certificates \
|
|
&& rm -rf /var/lib/apt/lists/*
|
|
|
|
# Pinned Python deps. Separate layer so source edits don't reinstall them.
|
|
COPY apps/pii/requirements.txt ./requirements.txt
|
|
RUN --mount=type=cache,target=/root/.cache/pip \
|
|
pip install -r requirements.txt
|
|
|
|
# Pinned spaCy models (en + es/it/pl/fi, ~2.2GB total). Downloaded with
|
|
# retries/resume — the large wheels truncate on flaky networks if pip fetches
|
|
# the URLs directly.
|
|
ARG SPACY_MODELS="en_core_web_lg-3.8.0 es_core_news_lg-3.8.0 it_core_news_lg-3.8.0 pl_core_news_lg-3.8.0 fi_core_news_lg-3.8.0"
|
|
RUN --mount=type=cache,target=/root/.cache/pip \
|
|
for model in ${SPACY_MODELS}; do \
|
|
whl="${model}-py3-none-any.whl"; \
|
|
curl -fL --retry 5 --retry-delay 5 --retry-all-errors -C - \
|
|
-o "/tmp/${whl}" \
|
|
"https://github.com/explosion/spacy-models/releases/download/${model}/${whl}" || exit 1; \
|
|
done && \
|
|
pip install /tmp/*.whl && \
|
|
rm /tmp/*.whl
|
|
|
|
COPY apps/pii/server.py ./server.py
|
|
|
|
RUN groupadd -g 1001 pii && \
|
|
useradd -u 1001 -g pii pii && \
|
|
chown -R pii:pii /app
|
|
USER pii
|
|
|
|
# Listen on 5001. Runs as its own ECS service (separate task), reached via PII_URL;
|
|
# 5001 avoids colliding with the app's 3000 in local/compose runs on one host.
|
|
EXPOSE 5001
|
|
|
|
# start-period is generous: five large spaCy models load at import before
|
|
# /health responds. Tune against measured cold-start once built.
|
|
HEALTHCHECK --interval=30s --timeout=5s --start-period=180s --retries=3 \
|
|
CMD curl -fsS http://localhost:5001/health || exit 1
|
|
|
|
CMD ["uvicorn", "server:app", "--host", "0.0.0.0", "--port", "5001"]
|