AI transparency — Guarrix
How Guarrix's detection systems are designed to work, and what they are not designed to do.
Guarrix is designed as an automated detection layer that inspects requests before they reach your LLM provider, and inspects responses before they reach your users. It is designed to identify personally identifiable information, prompt injection attempts, and jailbreak patterns, and to apply the action you configure — redact, mask, or block — plus any custom rules you define for your tenant.
Guarrix does not make decisions about people. It does not profile end users, does not generate content, and is not designed to produce any automated decision with legal or similarly significant effect on a person. It does not train on your prompt or response content — content is never stored.
Guarrix itself operates as a detection and filtering layer on your behalf, not as a system that makes autonomous decisions affecting people — it is not, in itself, a high-risk AI system under Annex III of the EU AI Act. Whether your own use of your chosen LLM is high-risk depends on how you use it; that classification is yours to make, not Guarrix's.
Article 50 of the EU AI Act requires that people be informed when they are interacting with an AI system, and that AI-generated or manipulated content be marked as such. Guarrix does not generate content shown to end users — it inspects and filters content already produced by your own application and LLM provider. If your own product uses AI to interact directly with end users, the Article 50 disclosure obligation towards those end users is yours to fulfil, not Guarrix's.
Guarrix's own detection engine is still in development and not yet in production use — this page will be updated with the specific detection models and providers once the service is live. Guarrix does not choose which LLM provider processes your customers' traffic; you configure your own provider (OpenAI, Anthropic, Mistral, or any OpenAI-compatible endpoint). Where Guarrix itself makes use of a large language model internally, it uses Mistral, hosted in the EU, for the same data-residency reasons as the rest of the Guarrix platform.
Guarrix is a probabilistic system. False positives (content incorrectly flagged) and false negatives (content incorrectly allowed through) are possible. Detection accuracy depends on the entity types and rules configured for your tenant. Guarrix is not designed or validated for use in safety-critical domains such as medical diagnosis, legal decisions, or any context where a missed detection could cause serious harm.
Questions about how Guarrix's AI systems are designed, or complaints about detection behaviour, can be sent to hello@guarrix.com. See our Privacy Policy for how to exercise your GDPR rights, and our Terms of Service for how disputes are handled.