Fast-moving teams shipping Gen AI features routinely introduce compliance hazards without meaning to: unmasked personal data (an SSN, a zip code) sitting in a prompt, race or gender quietly swaying a credit decision, an AI decision loop running with no one watching it.
These hazards hurt the people affected, and risk penalties under GDPR and the EU AI Act.
A small team can't afford a dedicated compliance hire, and reviewing every PR by hand would grind feature velocity to a crawl.
SHIP closes that gap: it catches AI compliance risk before it ships, without slowing the team down.
SHIP is a multi-agentic gatekeeper that reads every pull request the moment it opens, catches the ones that quietly cross a legal line, and interacts with you only when it's found something.
Ask it anytime through Alexa+. It answers in one sentence, and shows the full detail on your nearest screen.
Built for Build, Ship, Shape: the Amazon Developer Hackathon (Alexa+ track, AWS Builder mini-challenge).
SHIP is five pieces working as one pipeline, each named for its job: Screener, Detector, Triage, Gate, Relay.
- A PR opens, and Screener checks it first. A free regex/AST scan runs on every commit. Nothing matches, the PR passes in milliseconds, no model call spent.
- If something matches, Detector takes a closer look. It reviews just that fragment, grounded against sourced regulation text, not memory, and returns a verdict: risk score, plain-English explanation, citation, and a draft fix. It also learns from Gate's own past decisions over time.
- If Detector confirms a real violation, Triage blocks the merge. The PR's own commit status turns red, which is what actually stops the merge button, not just a dashboard entry.
- A human resolves it in Gate, the review console: file, risk, explanation, citation, suggested fix, one click to approve or reject. The decision posts back to the PR and clears the block.
- Relay carries that same review to voice and screen, so it comes to you wherever you are, instead of you having to go find it in Gate.
SHIP watches for four kinds of risk: exposing sensitive personal data, decisions made with no human review, bias from a protected trait, and AI handed more authority than it should have. Backed by six detectors tied to real law, full detail in architecture.html.
Three ways to try this.
- Experience it live. Judges: instructions for adding a test PR and trying it live are included in the submission materials.
- Watch the demo video. (Video link goes here before final submission.)
- Set it up yourself. Full steps in SETUP.md.
SHIP actively works in the background and alerts only when something needs your attention. Interact with it naturally, from your phone, an Echo, or anywhere else Alexa+ works.
1. Meets you where you are Findings aren't read aloud in a long monologue. Instead, anything that needs your review shows up on your nearest screen: a phone, a tablet, a laptop, a TV, even a smart fridge.
"Alexa, what's up?" "2 blockers found. Ready for your decision. Want to review now?" "Show me on the TV."
The actual findings then appear, live, on whatever screen just answered to that screen/device name.
2. Voice can confirm a decision. It can never make one blind. Ask to approve a finding you've never looked at, and SHIP refuses. That's the exact pattern SHIP is designed to flag in other people's code: an AI-mediated decision applied with no human checkpoint. Try it:
"Approve it." "Could you take a look on a screen first, then tell me why?"
Look at it first, then give a reason, and it goes through for real:
"Approve it. This is a false positive." "Done. Accepted, on the record: this is a false positive."
That gate is enforced server-side: voice only acts on a finding that's been shown on a screen in the last ten minutes, and a reason is required either way. That's a hard check, not a prompt asking the model to behave, so skip the review and there's nothing voice can say to talk its way past it.
Connecting a live Alexa+ device needs an account relationship set up by Amazon, not available during the hackathon. This uses a simulated Alexa+ experience instead, a web app calling the same deployed Relay backend, source included. Try it live.
1. Last line of defense. Assume nothing was checked before this. It doesn't matter who wrote the code, a person or an AI. It doesn't matter if the feature uses AI or not.
2. Grounded, not guessed. Rely on facts, not memory or guesswork. Keep improving to make that grounding sharper.
3. Meet you where you are. The answer finds you, on whatever device is nearest, not the other way around. Ask by voice from anywhere. See the detail on whatever screen is closest, a laptop, a TV, even a fridge.
4. Scalability and resilience. Survive retries, failures, and heavy load without breaking.
5. Cost efficiency. Free deterministic checks first, LLM only when needed. Only the right model for the job, measured.
Full diagram and component-by-component detail: pandayv.github.io/ship-engine (source).
The one piece worth calling out here: each flagged issue in a PR is reviewed as its own independent, retryable job, queued and picked up by an independently-scaling reviewer function, rather than one sequential pass through the whole PR. A fragment that fails outright retries automatically and lands in a dead-letter queue after repeated failure instead of vanishing. A fragment that's just slow, or that hits AWS's own request-rate limit, backs off and retries on its own without blocking anything else. A PR with several flagged issues takes about as long as its slowest single issue, not the sum of all of them, and nothing is silently dropped.
- Agent framework: Strands Agents SDK
- Models: Amazon Nova Lite as the primary Bedrock backend, chosen on
measured evidence rather than preference. Benchmarked against Claude
Haiku 4.5, Nova Pro, Qwen3-235B, GLM-5, and DeepSeek-V3.2 on the real
adversarial fixtures below: all six caught every genuine violation and
dismissed every look-alike, but per-model request-per-minute quota is
what actually bounds review throughput on this account (10/min for
Claude models vs. 200/min for Nova Lite). See the benchmark table in
src/agents/detector.py. Google Gemini serves as a credit-exhaustion fallback and a local Ollama model as a fully offline reliability fallback. Both are switchable via one env var. - Agent runtime: Amazon Bedrock AgentCore Runtime, the same Detector
logic deployed to a real managed runtime (
shipagentcore/), callable in-process for local development or remotely for the deployed path, toggled the same way - Retrieval: Amazon Bedrock Titan Embeddings, a small local numpy cosine-similarity store (the sourced regulation corpus is a few dozen chunks, so a hosted vector database would be pure overhead at this scale)
- Compute: two AWS Lambda functions, the webhook (fast-ack: verify, screen, dispatch) and an independently-scaling fragment processor (the actual model call, one fragment per invocation)
- Queueing: Amazon SQS, with a dead-letter queue for fragments that fail repeatedly and a concurrency cap on the processor so parallel reviews stay within the account's real request-rate limit
- State: Amazon DynamoDB, six tables.
ship-alerts(every write idempotent, so a webhook redelivery or a retried job can't create a duplicate and can't silently re-open a decision a human already made),ship-repos(which repositories SHIP watches, plus when each last sent an event and when one last finished review. The gap between the two, surfaced on Gate's connected-repos view, is what a silently broken pipeline looks like before anyone needs to notice by accident, seesrc/storage/repo_store.py),ship-status(a precomputed release summary, so Relay's voice path is oneGetItemregardless of how many findings exist, since Alexa+ allows Relay half a second to answer and aggregating on read would have blown that budget as findings accumulate),ship-device-connections(which device is reachable under which name, for the push path below),ship-recent-reviews(which alert was shown on a screen and when, the ten-minute windowrequest_risk_acceptancechecks before acting on a spoken confirmation), andship-learned-patterns(the generalized patterns above) - Web: FastAPI (the webhook route and the Gate console, one deployable app, wrapped for Lambda via Mangum)
- Voice/MCP: MCP Python SDK
(Streamable HTTP, spec
2025-11-25) powers Relay (src/relay/), deployed as its own Lambda with SnapStart enabled. A fresh ASGI app is built per invocation, a genuine SDK/Lambda incompatibility rather than a style choice. The docstring at the top ofsrc/relay/server.pyhas the details. - Push: an Amazon API Gateway WebSocket API plus a small dedicated
Lambda (
device_gateway_handler.py) handling connect/disconnect/register, deliberately separate from Relay so Relay's own IAM role stays scoped to exactly what answering a question requires. A display sits idle with no polling until Relay pushes a finding to it; a real compliance event happens on the order of weeks, not seconds, so a poll-every-few-seconds design would spend nearly all its traffic finding nothing changed
Deploying this yourself needs a real AWS account (Bedrock, Lambda,
DynamoDB, SQS, API Gateway) and about 30–45 minutes, less with
deploy.sh, which scripts most of it. Full step-by-step
walkthrough, including troubleshooting for a brand-new AWS account:
SETUP.md.
src/
agents/
screener.py # Fast regex/AST pre-filter + per-fragment isolation
detector.py # Strands Agent — RAG-grounded semantic judgment
triage.py # Risk-based routing, per-category thresholds
api/
main.py # Webhook route, fast-ack + fragment dispatch
dashboard.py # Gate — the human review console
github_client.py # Real PR-diff fetching via the GitHub API
github_writeback.py # Posts findings/decisions back to the PR (comment + commit status)
aws/
bedrock_session.py # Shared Bedrock session + adaptive-retry config
rag/
chunker.py # Regulation text -> citable paragraph chunks
vector_store.py # Local embedding store + retrieval
storage/
alert_store.py # DynamoDB — idempotent alert persistence
repo_store.py # DynamoDB — which repos SHIP watches
status_store.py # DynamoDB — precomputed release summary Relay's voice path reads
device_store.py # DynamoDB — which display device is reachable under which name
taxonomy.py # Single source of truth: detector <-> corpus <-> Screener trigger mapping
relay/
server.py # Relay — the MCP server; the voice/screen modality split lives here
modality.py # Enforces the spoken-response character cap and forbids line breaks
readmodel.py # Voice path (reads the precomputed summary) vs. screen path (reads alerts)
push.py # Delivers a payload to one named device over its open connection
rag_corpus/ # Sourced regulation/standard text, verbatim
shipagentcore/ # AgentCore Runtime deployment of Detector
docs/
alexa-simulator.html # The web-simulated Alexa+ experience — calls real, deployed Relay
ship-display.html # Any-device receiving surface — names itself, waits for a push
ship-render.js # Finding-rendering logic shared by both pages above
architecture.html # Full pipeline diagram
scripts/
healing_loop.py # Periodic corpus-grounding check, decoupled from the hot path
benchmark_models.py # The adversarial benchmark behind the Nova Lite model choice
precompute_embeddings.py # Regenerates rag_corpus/'s cached embeddings
lambda_handler.py # Webhook Lambda entrypoint
fragment_lambda_handler.py # Fragment-processor Lambda entrypoint
relay_lambda_handler.py # Relay's Lambda entrypoint — builds a fresh app per invocation, see server.py
device_gateway_handler.py # WebSocket connect/disconnect/register Lambda entrypoint
tests/ # 200 tests, no AWS credentials required to run
The full review pipeline, Screener through Gate, is live and deployed.
The PR's own merge button is gated by a commit status, not a dashboard
entry someone has to remember to check. Freezing an alert sets
ship/compliance to failing (promotable to a required check in branch
protection); resolving one recomputes it, closing the loop from detection
to a human decision to the merge button itself.
Relay (the MCP server), the Alexa+ simulated experience, and cross-device push are also live. See Integrated Alexa+ experience above for the full walkthrough.
What's not built yet:
- Auto-applying the suggested fix. Approving an alert doesn't push the patch to the PR yet; a human applies it after reviewing it in Gate. A GitHub-API integration away, not an architecture change.
- Scheduling the Healing Loop. Healing Loop re-checks that sourced regulation text still matches the law, catching anything a regulator has since amended. It runs today as a manual script, not yet on a schedule or alert channel.
- Connecting a live Alexa+ device. Needs an account relationship set up by Amazon, not available during the hackathon (see Integrated Alexa+ experience above). The simulator already exercises the same Relay endpoint and review logic; only the device transport is missing.
- Detectors that need more than a single PR diff. Infrastructure or deployment context, or behavior observed across multiple files or over time. See architecture.html for where that boundary sits.
MIT. See LICENSE.
