AI/ML & LLM security

AI Systems Make Decisions. Attackers Try to Influence Them.

LLM-powered assistants, agents and RAG pipelines introduce a new attack surface where untrusted text becomes instruction. Nuclisafe tests the application, model, tooling and retrieval layers together — because that is how they fail.

AI application attack surface

Every layer of an AI system is a testable boundary.

  1. 01

    User

    Malicious prompts, abuse, account takeover

  2. 02

    AI Application

    Prompt injection, weak auth, insecure output handling

  3. 03

    LLM / Model

    Jailbreaks, adversarial inputs, model abuse

  4. 04

    Tools / APIs

    Excessive agency, unauthorized actions

  5. 05

    Database / RAG

    Indirect injection, vector store leakage

  6. 06

    Business Systems

    Data exposure, privilege escalation

Data flows forward; trust does not. Untrusted content entering at any layer can influence decisions at every layer after it.

The AI nucleus: four surfaces, one system.

Prompt handling, model behaviour, retrieval sources and connected tools each carry distinct risks — and combine into chained attack paths. Testing them in isolation misses the exploits that matter.

What we test

Testing coverage across the attack surface

Coverage is tailored to your application. The areas below are assessed where applicable to the agreed scope.

Prompt InjectionIndirect Prompt InjectionJailbreak TestingSensitive Information DisclosureInsecure Output HandlingExcessive AgencyRAG SecurityVector Database SecurityModel / API SecurityAuthenticationAuthorizationData LeakageAdversarial InputsModel AbuseAI Supply Chain RisksData Poisoning Risks

Common weaknesses

Issues we frequently look for

Prompt Injection

User input overriding system instructions to change behaviour, leak context or reach restricted functionality.

Indirect Prompt Injection

Instructions hidden in retrieved documents, web pages or tool output that the model then obeys.

Sensitive Disclosure

System prompts, keys, internal documents or other users' data surfaced through model responses.

Insecure Output Handling

Model output rendered or executed downstream, enabling XSS, SSRF, SQL injection or command execution.

Excessive Agency

Agents with broad tool permissions performing unauthorized writes, purchases, emails or deletions.

RAG & Vector Store Exposure

Missing tenant isolation or document-level access control across embeddings and retrieval.

How we test

A manual-first testing approach

  • Architecture review of prompts, guardrails, tools, retrieval sources and trust boundaries.
  • Adversarial prompt testing: direct and indirect injection, jailbreaks, encoding and multi-turn manipulation.
  • Agent and tool abuse testing to establish what actions an attacker can trigger through the model.
  • Retrieval and vector store authorization testing across tenants, roles and document scopes.
  • Application-layer testing of the surrounding product: authentication, authorization, APIs and output rendering.

Aligned frameworks

OWASP Top 10 for LLM ApplicationsMITRE ATLASNIST AI RMFENISA AI Threat LandscapeGoogle SAIF

Assessments are mapped to these industry frameworks and testing methodologies. This does not imply certification by, or partnership with, any of these organisations.

Example tools

GarakIBM Adversarial Robustness ToolboxTextAttackFoolboxCleverHansMicrosoft PyRITLLM GuardRebuff

Illustrative only — tooling is selected per engagement and is not a guarantee of full coverage. Manual testing remains central to every assessment.

Deliverables

What you receive

Executive Summary

Business-level view of risk posture, key themes and priorities for leadership and stakeholders.

Technical Report

Detailed findings with affected endpoints, reproduction steps, evidence and references.

Proof of Concept

Validated demonstration of exploitability within the authorized scope, so nothing is theoretical.

Risk Rating

Severity based on impact and likelihood, supporting prioritization and remediation planning.

Remediation Guidance

Specific, actionable fix recommendations written for the developers who will implement them.

Retest Report

Post-fix verification confirming which findings are closed and which need further work.

Ready to start scoping?

Every assessment is scoped according to application complexity, attack surface and testing requirements.