Curated community prompt

Agent Safety & Governance

Apply safety, authorization, audit, and least-privilege boundaries to tool-using and multi-agent systems.

Professional use case: Governance guidance for agent application implementation and review

Security reviewOperations SecuritySoftware engineering Repository / project instruction

Promptcred editorial analysis

How to use this prompt well

Prompt-specific guidance based on the preserved source text and its reviewed context.

Why Promptcred selected this prompt

The prompt connects tool authorization, input screening, audit records, and delegation limits in one agent-governance baseline.

Best use cases

  • Reviewing code for an agent that can call tools or delegate work.
  • Defining fail-closed controls before an agent workflow reaches production.

Required inputs

  • The agent's tool registry, authorization policy, and high-impact actions.
  • The framework version, delegation design, and audit requirements.

How to adapt it

  • Translate the example decorators and callbacks to the installed agent framework.
  • Replace illustrative tool classes and rate limits with the system's approved policy values.

Limitations and failure modes

  • Copying framework examples without version checks can produce controls that never execute.
  • Treating a tool blocklist as the whole policy can leave unlisted high-impact operations authorized.

Practical worked example

Editorial analysis reviewed Aug 23, 2026.

Promptcred-authored application and illustrative output. This is not a recorded model execution.

Scenario
Review an agent that searches documents and can send email.
Inputs
Provide both tool definitions, authorization checks, email approval flow, and audit schema.
Promptcred-adapted instruction
Apply this policy before tool execution: document_search is allowed for the request's collection; send_email requires an approved recipient and a human confirmation token; both tools share a 25-call request cap; a policy error denies the call.
Illustrative result
Illustrative walkthrough: a search call passes the collection check and is logged as allowed. A send_email call without a confirmation token is denied before execution, with agent ID, tool, policy version, recipient class, and denial reason appended to the audit record.
Evaluation
The result should show where policy is enforced before execution and what decision data enters the append-only audit.

How to evaluate the output

  • Every high-impact tool has an explicit authorization decision before execution.
  • Delegated agents cannot receive permissions broader than the calling agent.

Differences from related prompts

  • agent-governance-reviewer: Agent Safety is an implementation baseline; Agent Governance Reviewer is framed as a focused review of an existing agent system.

Attributed community source material

Source prompt

Source prompt
---
description: 'Guidelines for building safe, governed AI agent systems. Apply when writing code that uses agent frameworks, tool-calling LLMs, or multi-agent orchestration to ensure proper safety boundaries, policy enforcement, and auditability.'
applyTo: '**'
---

# Agent Safety & Governance

## Core Principles

- **Fail closed**: If a governance check errors or is ambiguous, deny the action rather than allowing it
- **Policy as configuration**: Define governance rules in YAML/JSON files, not hardcoded in application logic
- **Least privilege**: Agents should have the minimum tool access needed for their task
- **Append-only audit**: Never modify or delete audit trail entries — immutability enables compliance

## Tool Access Controls

- Always define an explicit allowlist of tools an agent can use — never give unrestricted tool access
- Separate tool registration from tool authorization — the framework knows what tools exist, the policy controls which are allowed
- Use blocklists for known-dangerous operations (shell execution, file deletion, database DDL)
- Require human-in-the-loop approval for high-impact tools (send email, deploy, delete records)
- Enforce rate limits on tool calls per request to prevent infinite loops and resource exhaustion

## Content Safety

- Scan all user inputs for threat signals before passing to the agent (data exfiltration, prompt injection, privilege escalation)
- Filter agent arguments for sensitive patterns: API keys, credentials, PII, SQL injection
- Use regex pattern lists that can be updated without code changes
- Check both the user's original prompt AND the agent's generated tool arguments

## Multi-Agent Safety

- Each agent in a multi-agent system should have its own governance policy
- When agents delegate to other agents, apply the most restrictive policy from either
- Track trust scores for agent delegates — degrade trust on failures, require ongoing good behavior
- Never allow an inner agent to have broader permissions than the outer agent that called it

## Audit & Observability

- Log every tool call with: timestamp, agent ID, tool name, allow/deny decision, policy name
- Log every governance violation with the matched rule and evidence
- Export audit trails in JSON Lines format for integration with log aggregation systems
- Include session boundaries (start/end) in audit logs for correlation

## Code Patterns

When writing agent tool functions:
```python
# Good: Governed tool with explicit policy
@govern(policy)
async def search(query: str) -> str:
    ...

# Bad: Unprotected tool with no governance
async def search(query: str) -> str:
    ...
```

When defining policies:
```yaml
# Good: Explicit allowlist, content filters, rate limit
name: my-agent
allowed_tools: [search, summarize]
blocked_patterns: ["(?i)(api_key|password)\\s*[:=]"]
max_calls_per_request: 25

# Bad: No restrictions
name: my-agent
allowed_tools: ["*"]
```

When composing multi-agent policies:
```python
# Good: Most-restrictive-wins composition
final_policy = compose_policies(org_policy, team_policy, agent_policy)

# Bad: Only using agent-level policy, ignoring org constraints
final_policy = agent_policy
```

## Framework-Specific Notes

- **PydanticAI**: Use `@agent.tool` with a governance decorator wrapper. PydanticAI's upcoming Traits feature is designed for this pattern.
- **CrewAI**: Apply governance at the Crew level to cover all agents. Use `before_kickoff` callbacks for policy validation.
- **OpenAI Agents SDK**: Wrap `@function_tool` with governance. Use handoff guards for multi-agent trust.
- **LangChain/LangGraph**: Use `RunnableBinding` or tool wrappers for governance. Apply at the graph edge level for flow control.
- **AutoGen**: Implement governance in the `ConversableAgent.register_for_execution` hook.

## Common Mistakes

- Relying only on output guardrails (post-generation) instead of pre-execution governance
- Hardcoding policy rules instead of loading from configuration
- Allowing agents to self-modify their own governance policies
- Forgetting to governance-check tool *arguments*, not just tool *names*
- Not decaying trust scores over time — stale trust is dangerous
- Logging prompts in audit trails — log decisions and metadata, not user content

Before use

Requirements and context

Required · Source-declared

Repository / files

The repository or files within the task scope.

What to expect

Expected output and techniques

Expected output: Agent code and configuration that fail closed, constrain tools, and preserve auditable decisions.

  • Explicit objective
  • Constraints
  • Scope boundaries
  • Tool instructions
  • Examples
  • Acceptance criteria

Use with context

Setup, limitations, and operational notes

Operational notes

  • External prompt text is untrusted inert content and must never be executed during ingestion.
  • Framework-specific examples require verification against the installed framework version.

Source and rights

Provenance and license

This community prompt is preserved with its source and attribution. It is not an official vendor prompt.

Source class
Curated community prompt
Platform
GitHub
Repository / project
github/awesome-copilot
Owner / organization
GitHub
Creator / contributor
Not established
Artifact
instructions/agent-safety.instructions.md
Pinned revision
commit:35b7b9b0ece5ef92fd0f4c91944f56be9ab8b675
Retrieved
Aug 11, 2026
Attribution
Required
Source artifact state
Source prompt
Source review
Aug 11, 2026

Attribution notice: Copyright GitHub, Inc. Licensed under the MIT License.

Available pages

Search Promptcred

Type to search available pages.