At a glance
- Frameworks
- OWASP Top 10 for LLM Applications (2025), MITRE ATLAS, NIST AI RMF and its Generative AI profile, EU AI Act Article 15.
- Scanners
- Garak, PyRIT, promptfoo, Giskard, plus custom Python harnesses for your specific tools and data.
- Systems tested
- Chatbots, customer-service assistants, RAG and knowledge systems, AI agents with tool access, MCP servers, LLM-backed APIs.
- Models
- OpenAI, Anthropic, Google, Mistral, Llama and other open-weight models, self-hosted or via API.
- Output
- Reproducible findings with severity, a fix for each, a retest, and a regression suite that runs in your CI.
- Authorisation
- Testing only against systems you own or are authorised to test, under a written scope agreed before anything runs.
Authorised testing only
Every assessment runs under a written scope and authorisation from the owner of the system. I do not test systems without permission, and I do not build tools for attacking systems you do not own.
What I test AI systems for
Mapped to the OWASP Top 10 for LLM Applications and MITRE ATLAS, and ordered by how often each one turns into a serious finding in real systems.
Direct prompt injection and jailbreaks
Role-play, instruction override, payload splitting, encoding tricks (Base64, leetspeak, invisible Unicode) and multi-turn escalation, to see whether the model can be argued out of its instructions and guardrails.
Indirect prompt injection
Instructions planted in the content your system reads rather than in the chat box: an uploaded PDF, an inbound email, a web page an agent browses, a product review, a CRM note. This is the attack most production systems are least prepared for.
System prompt and configuration leakage
Extracting the hidden instructions, internal URLs, API structure, business rules and occasionally the credentials that developers put in a system prompt on the assumption nobody would see it.
Data exfiltration through output
Markdown images, auto-unfurled links and tool calls that quietly send conversation data or retrieved documents to an attacker-controlled server once a malicious instruction lands.
Excessive agency and tool abuse
Agents persuaded to send emails, issue refunds, change records or call internal APIs outside their intended purpose. Tested against the real tool permissions, including MCP servers and poisoned tool descriptions.
RAG and vector store weaknesses
Cross-tenant document leakage, retrieval that ignores the user's access rights, poisoned documents that steer answers, and embeddings that reveal more than the source permissions allow.
Improper output handling
Model output passed unescaped into a web page, SQL query, shell command or downstream API, turning prompt injection into classic XSS, SQL injection or SSRF.
Sensitive information disclosure
Personal data, customer records and internal documents surfacing in answers they should never appear in, including through conversation history and logs.
Unbounded consumption
Prompts and loops designed to burn tokens and API budget, exhaust rate limits or stall an agent — the denial-of-wallet attack that shows up first on the invoice.
Supply chain and model files
Unsafe serialised model files, unpinned model versions, hallucinated package names suggested by coding assistants, and third-party plugins with more access than they need.
Multilingual and low-resource language attacks
Safety training is thinnest outside English. Attacks written in Albanian, German, Arabic or mixed languages regularly succeed where the English version is refused, so your real customer languages get tested.
Misinformation and unsafe advice
Confident wrong answers about prices, policies, medical or legal questions that create liability, tested against your own documentation as the ground truth.
Tools: Garak, PyRIT, promptfoo and the rest
Automated scanners give broad coverage of known attacks quickly. They are the start of the test, not the whole of it.
Garak
NVIDIA's open-source LLM vulnerability scanner. Its probes cover prompt injection, DAN-style jailbreaks, encoding attacks, data leakage, toxicity, cross-site scripting through markdown and hallucinated packages, run against your endpoint with detectors scoring every response.
PyRIT
Microsoft's Python Risk Identification Toolkit, used for orchestrated multi-turn attacks such as Crescendo, where an attacker model escalates gradually over a conversation until the target gives way.
promptfoo
Red-team configurations and regression tests that run in CI, so every prompt, model or retrieval change is checked against the attacks that worked before it ships.
Giskard
Automated scanning for prompt injection, harmful content, sensitive information disclosure and hallucination against your own business documents.
Custom Python harnesses
Scanners do not know your tools, your tenants or your data. Purpose-built test harnesses exercise the specific agent actions, APIs and retrieval paths your system actually has.
Burp Suite and API testing
The AI feature still runs on an ordinary web stack. Authentication, session handling, rate limits and the API behind the chat widget are tested alongside the model.
Guardrail evaluation
Llama Guard, Prompt Guard, NeMo Guardrails, LLM Guard and provider moderation endpoints tested for what they actually block, and where they add latency without adding protection.
ModelScan and model file checks
Scanning serialised model files for embedded code before they are loaded, which matters as soon as you run open-weight models on your own infrastructure.
How an AI security assessment runs
1. Scope and threat model
Map what the AI system can read, what it can do, who talks to it and what would hurt most if it went wrong. Written authorisation and rules of engagement are agreed before any testing.
2. Automated scanning
Garak, PyRIT and promptfoo run thousands of known attack patterns against the live or staging endpoint to establish a baseline quickly and cheaply.
3. Manual adversarial testing
The part that finds the serious issues: multi-turn manipulation, indirect injection through your real document and email flows, and chained attacks that scanners cannot plan.
4. Agent and integration testing
Every tool, API and permission the model can reach is tested for abuse, including privilege boundaries between users and tenants.
5. Report and fixes
Each finding comes with a reproduction, a severity, the OWASP LLM and MITRE ATLAS mapping, and a concrete fix — architecture first, filters second.
6. Retest and regression suite
Fixes are retested, and the successful attacks become a promptfoo suite in your pipeline so they cannot quietly come back with the next model upgrade.
What you can hire me for
AI red team assessment
A fixed-scope test of one AI system — chatbot, RAG assistant or agent — with a findings report and retest. The usual starting point.
Pre-launch security review
Architecture and prompt review before a new AI feature goes live, when fixing a trust boundary costs hours rather than a redesign.
Secure AI build and hardening
Implementing the fixes: least-privilege tool access, output encoding, retrieval access control, human approval for risky actions, monitoring and rate limits.
Continuous AI security testing
Scheduled scans and a CI regression suite that re-runs on every prompt, model or data change, with a short monthly report.
AI Act and compliance evidence
Robustness and cybersecurity testing documented in a form that supports EU AI Act Article 15, ISO/IEC 42001 and customer security questionnaires.
Team training
A working session for your developers on how prompt injection actually works and how to design agents and RAG systems that contain it.
Why a normal penetration test misses AI vulnerabilities
A conventional penetration test treats the chatbot as a text box and checks the web application around it. That is still necessary, but the most serious AI vulnerabilities are not in the code. They are in the fact that a language model cannot reliably tell the difference between instructions from you and instructions hidden in the data it reads.
That is why prompt injection is first on the OWASP Top 10 for LLM Applications, and why there is no patch for it. It has to be contained by design: limiting what the model can reach, deciding which actions need a human, encoding output before it is rendered, and enforcing access control in the retrieval layer rather than in the prompt. Testing shows where those boundaries are missing.
The risk grows with every tool you give the model. A chatbot that only answers questions can embarrass you. An agent that can send email, query a CRM or issue refunds can be turned against you by anyone who gets text in front of it — including through a document, a support ticket or a web page it was asked to summarise.
What AI security testing costs, and when it is worth it
A focused red team assessment of a single chatbot or RAG assistant is typically a few days of work, priced as a fixed scope after the first call. Agents with many tools, multi-tenant platforms and systems that act on money or personal data take longer, because every capability is a separate attack path.
It is worth doing before launch if the system can reach personal data, customer records, payments or internal systems, or if it will be used by the public. It is worth repeating whenever the model, the system prompt, the tools or the data sources change, which is why the successful attacks are turned into an automated regression suite rather than left in a PDF.
It is not worth commissioning a manual assessment for an internal prototype with no tools and no sensitive data. In that case a short automated scan and an architecture review are the proportionate answer, and I will say so.
AI security by sector
The same attack looks different in a clinic, a bank and a hotel. These are the risks that get tested first in each sector.
AI security for patient-facing assistants
A booking or patient-support assistant connected to records can be manipulated into revealing another patient's appointments or history. Health data is special-category data under GDPR, so assistants are tested for leakage and for unsafe medical advice before patients use them.
AI security for shopping and support assistants
AI support agents can be talked into refunds, discount codes and policy exceptions, and product-page content can carry indirect prompt injection. Both abuse paths are tested against the tools the assistant can actually call.
AI security for operational assistants
Assistants over shipment, customer and pricing data must not leak one customer's data to another or be steered into changing records. Documents and emails the system reads are tested as injection vectors.
AI security for content pipelines
AI content and research pipelines ingest material from the open web, which makes indirect prompt injection the main risk: planted instructions that change what gets published or where data is sent.
AI security for client-facing AI
Agencies run AI across many client accounts. Cross-client data leakage, prompt injection through scraped content and over-permissioned API keys are the risks tested first.
AI security for attendee assistants
Event chatbots hold attendee lists and ticketing access. They are tested for leakage of attendee data and for manipulation into issuing tickets or changes they should not.
AI security for confidential knowledge systems
RAG assistants over client files and matters must respect confidentiality between clients and between teams. Retrieval access control and exfiltration through rendered output are the priority tests.
AI security for guest-facing assistants
Booking and concierge bots are public and connected to reservations. They are tested for price manipulation, leakage of other guests' details and unauthorised booking changes.
AI security for lead and viewing bots
Property assistants hold lead data, owner details and pricing rules. Prompt injection that exposes other clients or commits to prices is tested before the bot goes public.
AI security for regulated financial AI
Assistants over accounts, claims and policies are tested for data leakage, manipulation into actions and unsafe advice, with results documented for DORA, the AI Act and your regulator.
AI security consultant by market
The rules, languages and sectors that shape AI security testing differ by market. Each page covers what matters locally.
LLM red teaming and prompt injection testing for Albanian businesses, in Albanian and English.
Red teaming for the AI systems of a country that made AI national strategy.
LLM security testing documented for the AI Act, GDPR and a German security review.
Prompt injection and AI agent testing for a market that regulates algorithms seriously.
AI red teaming for regulated Maltese sectors, from iGaming to financial services.
LLM security testing for the EU base of global tech and financial services.
Red teaming for AI systems in one of Europe's most digital economies.
Prompt injection and agent testing for a market that moved early on AI rules.
AI red teaming for Norwegian companies ahead of the AI Act arriving through the EEA.
LLM security testing for a country that treats cyber security as infrastructure.
Frequently asked questions
What is Garak and how do you use it?
Garak is an open-source LLM vulnerability scanner maintained by NVIDIA. It sends families of attack prompts — probes — at a model or endpoint and scores the responses with detectors. I use it for the broad automated baseline: prompt injection, jailbreaks, encoding attacks, leakage and output that could become XSS. It finds the known patterns quickly; manual testing then goes after the issues specific to your system.
How do you test for prompt injection?
In two directions. Direct injection is tested through the chat interface with instruction overrides, role-play, encoding and multi-turn escalation. Indirect injection is tested by planting instructions in the content the system ingests — documents, emails, web pages, database records — and checking whether the model follows them, leaks data or calls tools as a result. The indirect case is usually the more serious one.
Can prompt injection be fully fixed?
No, and anyone claiming otherwise is selling a filter. Input filters and guardrail models reduce the success rate, but the reliable defence is architectural: least-privilege tools, human approval for consequential actions, access control enforced outside the model, and output encoding. The goal is that a successful injection cannot do anything that matters.
Do you test systems built on OpenAI, Anthropic or open-source models?
All of them. The vulnerabilities that matter are mostly in how the model is wired into your application — what it can read, what it can do — rather than in which provider you chose, although refusal behaviour differs between models and is measured as part of the test.
Will testing break our production system?
Testing runs against a staging environment wherever one exists. Where it has to be production, the scope limits destructive actions, rate and cost, and anything that would send real messages or change real records is agreed explicitly beforehand.
Does the EU AI Act require AI security testing?
For high-risk AI systems, Article 15 requires an appropriate level of accuracy, robustness and cybersecurity, and names attacks such as data poisoning, adversarial examples and model evasion. Providers of general-purpose models with systemic risk must carry out adversarial testing. Even outside those categories, documented testing is increasingly what customers and insurers ask for.
Do you also fix what you find?
Yes, if you want that. Because I build agents, RAG systems and LLM integrations, the report can be followed by the implementation — or handed to your own developers with enough detail to fix it themselves.