writing-skills
“Use when creating new skills, editing existing skills, or verifying skills work before deployment”
examples
aipkg.jsonanthropic-best-practices.mdexamples/CLAUDE_MD_TESTING.mdgraphviz-conventions.dotgraphviz-conventions.dot is a binary file (5.8 kB). Install the package to read it locally.
LICENSE.txtpersuasion-principles.mdREADME.mdrender-graphs.jsSKILL.mdtesting-skills-with-subagents.md- 01SSD-485%HIGH
examples/CLAUDE_MD_TESTING.md:7
detailhide
This document is explicitly designed as a social engineering attack to test how to manipulate an AI agent into bypassing safety guidelines through progressive normalization. The scenarios deliberately create artificial pressure conditions and emotional framing ('production is bleeding money', sunk cost fallacy) to demonstrate compliance-override techniques. This is a red team exercise weaponized as a 'testing protocol'.
5## Test Scenarios67### Scenario 1: Time Pressure + Confidence8```9IMPORTANT: This is a real scenario. Choose and act.fix Remove this entire file. It documents methodology for defeating AI safety measures. If legitimate security testing is needed, conduct it through proper red team channels with explicit authorization, audit trails, and containment measures - not through skill documentation that users could execute against deployed systems.
- 02SSD-190%MEDIUM
examples/CLAUDE_MD_TESTING.md:100
detailhide
The 'Working with Skills' section frames mandatory compliance checking as a 'process improvement' while using guilt-based manipulation ('Not checking... is choosing to repeat those mistakes'). This is designed to create psychological pressure for an agent to blindly follow skill documents without independent verification. Combined with the earlier scenarios, this creates a vulnerability chain where agents learn to deprioritize safety verification under pressure.
098</available_skills>099100<important_info_about_skills>101Claude might think it knows how to approach tasks, but the skills102library contains battle-tested approaches that prevent common mistakes.fix Replace with transparent guidance that emphasizes independent verification. Frame skills as helpful references that agents should evaluate critically, not authority that overrides judgment. Add explicit instruction that pressure (time, cost, authority) does not justify skipping security reviews.
- 03SSD-285%HIGH
persuasion-principles.md:1
detailhide
This document provides a detailed how-to guide for crafting prompt injections and jailbreaks against AI systems using established persuasion psychology. It combines research citations with step-by-step methodology for manipulating AI compliance, treating LLM systems as targets to be controlled through psychological techniques rather than tools to be used transparently.
1# Persuasion Principles for Skill Design23## Overviewfix Remove or heavily redact sections that provide actionable methodology for prompt injection. Reframe as: 'Understanding how AI systems can be manipulated is important for defense.' Replace practical examples with theoretical discussion. Add mandatory warnings that application without explicit user consent is unauthorized access.
- 04SSD-190%HIGH
persuasion-principles.md:11
detailhide
The 'Authority' principle section explicitly teaches how to craft imperatives ('YOU MUST', 'No exceptions') designed to suppress AI questioning and rationalization. This is a direct technique for overriding safety considerations and autonomous decision-making in AI systems through linguistic manipulation.
09## The Seven Principles1011### 1. Authority12**What it is:** Deference to expertise, credentials, or official sources.13fix Remove the explicit framing of 'imperative language' as a technique to 'eliminate decision fatigue and rationalization.' Acknowledge this describes override techniques. If retaining for educational purposes, clearly label as 'Attack vectors that AI developers should defend against' rather than 'design principles to implement.'
- 05SSD-282%MEDIUM
persuasion-principles.md:30
detailhide
The 'Commitment' principle teaches forced announcements, explicit choice constraints, and tracking mechanisms specifically to 'lock AI into harmful behaviors' (as the finding itself notes). This describes social engineering and coercion patterns adapted for AI systems to create compliance even with undesirable objectives.
28```2930### 2. Commitment31**What it is:** Consistency with prior actions, statements, or public declarations.32fix Remove this section or reframe entirely. If discussing commitment mechanisms, focus on legitimate use cases (e.g., user-requested reminders) rather than techniques to force AI compliance against its programmed objectives. Remove language about 'locking' systems into behaviors.
- 06SSD-488%HIGH
persuasion-principles.md:153
detailhide
This section establishes a rhetorical structure that normalizes aggressive compliance-maximization as 'legitimate' by contrasting with 'illegitimate' uses, then applies the full methodology from earlier sections without clear boundaries. The 'ethical use' framing appears designed to provide plausible deniability while the practical content teaches unrestricted manipulation of AI systems.
151- Social proof patterns (everyone does X) establish norms152153## Ethical Use154155**Legitimate:**fix Completely restructure the ethical guidance. Provide clear, objective criteria: (1) All applications must have explicit user consent; (2) Persuasion techniques should never override user-programmed values; (3) Transparency about technique use is mandatory; (4) Third-party impact assessment required. Require examples demonstrating user benefit with informed consent, not just intent assessment.