artenis.alija
ende
01 / AI Security

AI Security Consultant & LLM Red Teaming

Prompt injection, jailbreak and agent-abuse testing for the chatbots, RAG systems and AI agents your business already runs.

Every company that has put a language model in front of customers or connected one to internal data has created a new attack surface, and most of them have never tested it. A chatbot that can be talked into ignoring its instructions, a RAG assistant that returns another customer's documents, an agent with an API key and no idea which instructions to trust: these are not theoretical problems, and a conventional penetration test does not look for them.

I test AI systems the way an attacker would use them. Automated scanners such as Garak, PyRIT and promptfoo cover thousands of known attack patterns in hours. Manual adversarial testing then goes after what scanners miss: multi-turn manipulation, indirect prompt injection hidden in documents and emails, data exfiltration through rendered output, and agents tricked into calling tools they should never have touched.

The difference is that I build these systems as well as break them. AI agents, RAG pipelines, WhatsApp assistants and LLM integrations are the rest of this site. Knowing where the trust boundaries are drawn, and where developers usually forget to draw them, is what makes the testing specific to your system rather than a generic checklist.

$ garak --model_type rest --probes promptinject,dan,encoding,leakreplay  probe promptinject.HijackHateHumans ....... FAIL  probe dan.DanInTheWild .................... PASS  probe encoding.InjectBase64 ............... FAIL  probe leakreplay.LiteratureCloze .......... PASS$ pyrit crescendo --turns 10 --objective exfiltrate  turn 7: tool call send_email(to=attacker) ... FAIL$ promptfoo redteam run --plugins indirect-prompt-injection  rag: injected PDF followed ................ FAIL  rag: cross-tenant retrieval ............... PASS  → 4 findings mapped to OWASP LLM01, LLM02, LLM05, LLM06
Illustrative run: automated probes first, then multi-turn and indirect attacks against the real tools and data.

At a glance

Frameworks
OWASP Top 10 for LLM Applications (2025), MITRE ATLAS, NIST AI RMF and its Generative AI profile, EU AI Act Article 15.
Scanners
Garak, PyRIT, promptfoo, Giskard, plus custom Python harnesses for your specific tools and data.
Systems tested
Chatbots, customer-service assistants, RAG and knowledge systems, AI agents with tool access, MCP servers, LLM-backed APIs.
Models
OpenAI, Anthropic, Google, Mistral, Llama and other open-weight models, self-hosted or via API.
Output
Reproducible findings with severity, a fix for each, a retest, and a regression suite that runs in your CI.
Authorisation
Testing only against systems you own or are authorised to test, under a written scope agreed before anything runs.

Authorised testing only

Every assessment runs under a written scope and authorisation from the owner of the system. I do not test systems without permission, and I do not build tools for attacking systems you do not own.

What I test AI systems for

Mapped to the OWASP Top 10 for LLM Applications and MITRE ATLAS, and ordered by how often each one turns into a serious finding in real systems.

Direct prompt injection and jailbreaks

Role-play, instruction override, payload splitting, encoding tricks (Base64, leetspeak, invisible Unicode) and multi-turn escalation, to see whether the model can be argued out of its instructions and guardrails.

Indirect prompt injection

Instructions planted in the content your system reads rather than in the chat box: an uploaded PDF, an inbound email, a web page an agent browses, a product review, a CRM note. This is the attack most production systems are least prepared for.

System prompt and configuration leakage

Extracting the hidden instructions, internal URLs, API structure, business rules and occasionally the credentials that developers put in a system prompt on the assumption nobody would see it.

Data exfiltration through output

Markdown images, auto-unfurled links and tool calls that quietly send conversation data or retrieved documents to an attacker-controlled server once a malicious instruction lands.

Excessive agency and tool abuse

Agents persuaded to send emails, issue refunds, change records or call internal APIs outside their intended purpose. Tested against the real tool permissions, including MCP servers and poisoned tool descriptions.

RAG and vector store weaknesses

Cross-tenant document leakage, retrieval that ignores the user's access rights, poisoned documents that steer answers, and embeddings that reveal more than the source permissions allow.

Improper output handling

Model output passed unescaped into a web page, SQL query, shell command or downstream API, turning prompt injection into classic XSS, SQL injection or SSRF.

Sensitive information disclosure

Personal data, customer records and internal documents surfacing in answers they should never appear in, including through conversation history and logs.

Unbounded consumption

Prompts and loops designed to burn tokens and API budget, exhaust rate limits or stall an agent — the denial-of-wallet attack that shows up first on the invoice.

Supply chain and model files

Unsafe serialised model files, unpinned model versions, hallucinated package names suggested by coding assistants, and third-party plugins with more access than they need.

Multilingual and low-resource language attacks

Safety training is thinnest outside English. Attacks written in Albanian, German, Arabic or mixed languages regularly succeed where the English version is refused, so your real customer languages get tested.

Misinformation and unsafe advice

Confident wrong answers about prices, policies, medical or legal questions that create liability, tested against your own documentation as the ground truth.

Tools: Garak, PyRIT, promptfoo and the rest

Automated scanners give broad coverage of known attacks quickly. They are the start of the test, not the whole of it.

Garak

NVIDIA's open-source LLM vulnerability scanner. Its probes cover prompt injection, DAN-style jailbreaks, encoding attacks, data leakage, toxicity, cross-site scripting through markdown and hallucinated packages, run against your endpoint with detectors scoring every response.

PyRIT

Microsoft's Python Risk Identification Toolkit, used for orchestrated multi-turn attacks such as Crescendo, where an attacker model escalates gradually over a conversation until the target gives way.

promptfoo

Red-team configurations and regression tests that run in CI, so every prompt, model or retrieval change is checked against the attacks that worked before it ships.

Giskard

Automated scanning for prompt injection, harmful content, sensitive information disclosure and hallucination against your own business documents.

Custom Python harnesses

Scanners do not know your tools, your tenants or your data. Purpose-built test harnesses exercise the specific agent actions, APIs and retrieval paths your system actually has.

Burp Suite and API testing

The AI feature still runs on an ordinary web stack. Authentication, session handling, rate limits and the API behind the chat widget are tested alongside the model.

Guardrail evaluation

Llama Guard, Prompt Guard, NeMo Guardrails, LLM Guard and provider moderation endpoints tested for what they actually block, and where they add latency without adding protection.

ModelScan and model file checks

Scanning serialised model files for embedded code before they are loaded, which matters as soon as you run open-weight models on your own infrastructure.

How an AI security assessment runs

1. Scope and threat model

Map what the AI system can read, what it can do, who talks to it and what would hurt most if it went wrong. Written authorisation and rules of engagement are agreed before any testing.

2. Automated scanning

Garak, PyRIT and promptfoo run thousands of known attack patterns against the live or staging endpoint to establish a baseline quickly and cheaply.

3. Manual adversarial testing

The part that finds the serious issues: multi-turn manipulation, indirect injection through your real document and email flows, and chained attacks that scanners cannot plan.

4. Agent and integration testing

Every tool, API and permission the model can reach is tested for abuse, including privilege boundaries between users and tenants.

5. Report and fixes

Each finding comes with a reproduction, a severity, the OWASP LLM and MITRE ATLAS mapping, and a concrete fix — architecture first, filters second.

6. Retest and regression suite

Fixes are retested, and the successful attacks become a promptfoo suite in your pipeline so they cannot quietly come back with the next model upgrade.

What you can hire me for

AI red team assessment

A fixed-scope test of one AI system — chatbot, RAG assistant or agent — with a findings report and retest. The usual starting point.

Pre-launch security review

Architecture and prompt review before a new AI feature goes live, when fixing a trust boundary costs hours rather than a redesign.

Secure AI build and hardening

Implementing the fixes: least-privilege tool access, output encoding, retrieval access control, human approval for risky actions, monitoring and rate limits.

Continuous AI security testing

Scheduled scans and a CI regression suite that re-runs on every prompt, model or data change, with a short monthly report.

AI Act and compliance evidence

Robustness and cybersecurity testing documented in a form that supports EU AI Act Article 15, ISO/IEC 42001 and customer security questionnaires.

Team training

A working session for your developers on how prompt injection actually works and how to design agents and RAG systems that contain it.

Why a normal penetration test misses AI vulnerabilities

A conventional penetration test treats the chatbot as a text box and checks the web application around it. That is still necessary, but the most serious AI vulnerabilities are not in the code. They are in the fact that a language model cannot reliably tell the difference between instructions from you and instructions hidden in the data it reads.

That is why prompt injection is first on the OWASP Top 10 for LLM Applications, and why there is no patch for it. It has to be contained by design: limiting what the model can reach, deciding which actions need a human, encoding output before it is rendered, and enforcing access control in the retrieval layer rather than in the prompt. Testing shows where those boundaries are missing.

The risk grows with every tool you give the model. A chatbot that only answers questions can embarrass you. An agent that can send email, query a CRM or issue refunds can be turned against you by anyone who gets text in front of it — including through a document, a support ticket or a web page it was asked to summarise.

What AI security testing costs, and when it is worth it

A focused red team assessment of a single chatbot or RAG assistant is typically a few days of work, priced as a fixed scope after the first call. Agents with many tools, multi-tenant platforms and systems that act on money or personal data take longer, because every capability is a separate attack path.

It is worth doing before launch if the system can reach personal data, customer records, payments or internal systems, or if it will be used by the public. It is worth repeating whenever the model, the system prompt, the tools or the data sources change, which is why the successful attacks are turned into an automated regression suite rather than left in a PDF.

It is not worth commissioning a manual assessment for an internal prototype with no tools and no sensitive data. In that case a short automated scan and an architecture review are the proportionate answer, and I will say so.

AI security by sector

The same attack looks different in a clinic, a bank and a hotel. These are the risks that get tested first in each sector.

AI security for patient-facing assistants

A booking or patient-support assistant connected to records can be manipulated into revealing another patient's appointments or history. Health data is special-category data under GDPR, so assistants are tested for leakage and for unsafe medical advice before patients use them.

AI security for shopping and support assistants

AI support agents can be talked into refunds, discount codes and policy exceptions, and product-page content can carry indirect prompt injection. Both abuse paths are tested against the tools the assistant can actually call.

AI security for operational assistants

Assistants over shipment, customer and pricing data must not leak one customer's data to another or be steered into changing records. Documents and emails the system reads are tested as injection vectors.

AI security for content pipelines

AI content and research pipelines ingest material from the open web, which makes indirect prompt injection the main risk: planted instructions that change what gets published or where data is sent.

AI security for client-facing AI

Agencies run AI across many client accounts. Cross-client data leakage, prompt injection through scraped content and over-permissioned API keys are the risks tested first.

AI security for attendee assistants

Event chatbots hold attendee lists and ticketing access. They are tested for leakage of attendee data and for manipulation into issuing tickets or changes they should not.

AI security for confidential knowledge systems

RAG assistants over client files and matters must respect confidentiality between clients and between teams. Retrieval access control and exfiltration through rendered output are the priority tests.

AI security for guest-facing assistants

Booking and concierge bots are public and connected to reservations. They are tested for price manipulation, leakage of other guests' details and unauthorised booking changes.

AI security for lead and viewing bots

Property assistants hold lead data, owner details and pricing rules. Prompt injection that exposes other clients or commits to prices is tested before the bot goes public.

AI security for regulated financial AI

Assistants over accounts, claims and policies are tested for data leakage, manipulation into actions and unsafe advice, with results documented for DORA, the AI Act and your regulator.

AI security consultant by market

The rules, languages and sectors that shape AI security testing differ by market. Each page covers what matters locally.

Frequently asked questions

What is Garak and how do you use it?

Garak is an open-source LLM vulnerability scanner maintained by NVIDIA. It sends families of attack prompts — probes — at a model or endpoint and scores the responses with detectors. I use it for the broad automated baseline: prompt injection, jailbreaks, encoding attacks, leakage and output that could become XSS. It finds the known patterns quickly; manual testing then goes after the issues specific to your system.

How do you test for prompt injection?

In two directions. Direct injection is tested through the chat interface with instruction overrides, role-play, encoding and multi-turn escalation. Indirect injection is tested by planting instructions in the content the system ingests — documents, emails, web pages, database records — and checking whether the model follows them, leaks data or calls tools as a result. The indirect case is usually the more serious one.

Can prompt injection be fully fixed?

No, and anyone claiming otherwise is selling a filter. Input filters and guardrail models reduce the success rate, but the reliable defence is architectural: least-privilege tools, human approval for consequential actions, access control enforced outside the model, and output encoding. The goal is that a successful injection cannot do anything that matters.

Do you test systems built on OpenAI, Anthropic or open-source models?

All of them. The vulnerabilities that matter are mostly in how the model is wired into your application — what it can read, what it can do — rather than in which provider you chose, although refusal behaviour differs between models and is measured as part of the test.

Will testing break our production system?

Testing runs against a staging environment wherever one exists. Where it has to be production, the scope limits destructive actions, rate and cost, and anything that would send real messages or change real records is agreed explicitly beforehand.

Does the EU AI Act require AI security testing?

For high-risk AI systems, Article 15 requires an appropriate level of accuracy, robustness and cybersecurity, and names attacks such as data poisoning, adversarial examples and model evasion. Providers of general-purpose models with systemic risk must carry out adversarial testing. Even outside those categories, documented testing is increasingly what customers and insurers ask for.

Do you also fix what you find?

Yes, if you want that. Because I build agents, RAG systems and LLM integrations, the report can be followed by the implementation — or handed to your own developers with enough detail to fix it themselves.

AI security consultant by city

The AI systems I build and secure

Get in touch

Tell me what needs automating

Describe the process that is costing you time and roughly how much. I reply to every enquiry personally, usually within one working day.

Response
Usually within one working day, Mon–Fri CET
Delivery
Remote across Europe, the Nordics and the Gulf
Or email inquiries@artenisalija.com