Beyond AI Ethics: The Rise of the Guardrail Angels
Anna AugustShare
The Unattended Space in AI Ethics
Yesterday I listened to several podcasts and interviews with leading AI thinkers (after all, Sunday is a day for inspiration). It left me reflecting on the gaps that still exist in AI ethics—areas that remain largely unattended, where a human hand is still needed to guide how AI can collaborate more effectively with people.
At present, our interaction with AI relies primarily on the written word. The prompt is the foundation of what we do. In the beginning was the word. Yet not all words are equal. What LLMs specialise in - tokenisation and predicting the next symbol based on massive datasets - is, at its core, a purely mathematical construct: derivative, dry, and statistical. This can generate a range of problems, beginning with bias—the reproduction of what is most represented in the data, which is not necessarily what is true, but simply what appears most often.
Pure Probability vs. Social Context
To simplify: if I begin a word with “FU…,” and “fuck” happens to be the most frequent continuation in the dataset, a model would be designed to output exactly that. Pure mathematics. This is roughly how early neural networks (RNNs) behaved. Today, however, LLMs are layered with additional mechanisms—context-sensitive probability distributions, randomness parameters, safety filters, and other guardrails—introducing forms of control over generated content. Because the word in question is considered offensive, it may be pushed down the probability ranking and replaced by something like “function.” By default, according to pure statistical design, “fuck” might have appeared; in practice, it is contextualised, deprioritised, and filtered according to social norms.
The Complexity of Human Language
Language, however, is extraordinarily complex—full of nuance, exceptions, ambiguities, and cultural signals. It is therefore impossible to govern linguistic behaviour solely through general rules. Consider Polish, a language often described as one of the most difficult to learn precisely because it is built on layers of exceptions. Native speakers, having grown up immersed in it, can use it effortlessly, like virtuosos playing an instrument after years of practice—often without consciously knowing the grammatical rules that underpin their fluency. Yet their communicative precision remains unquestionable.
The same is true of prompts. If we feed machines only symbols without context, the result is a dry mathematical artefact—sometimes accurate, sometimes flawed, and at times hallucinatory. Context is everything. A human raised within a language intuitively understands its usage; likewise, a machine trained to recognise contextual patterns can learn to communicate appropriately.
This is where AI ethicists enter, though the title does not quite capture the role. A more fitting name might be Prompt Editor, Cultural Guardrail Analyst, or even Guardrail Angel. Nor am I referring to the role that has begun appearing on the market in 2026 under the name AI Language Specialist. What I have in mind reaches well beyond the domain of a mere language expert.
Translating Culture for Machines
Such individuals would be responsible for identifying context, especially cultural and communicative context: what makes sense, what sounds natural, what feels awkward or inappropriate. They would ensure that AI systems operate within appropriate social norms, understand irony and provocation, and recognise emotional registers such as anger, joy, or satisfaction. They would understand that what is considered normal in one country may provoke discomfort, conflict, or fear in another. A Guardrail Angel tests and calibrates systems in domains prone to escalation—politics, religion, violence, stereotypes—not to censor uncomfortable topics, but to ensure that responses are thoughtful rather than mechanical (“That’s a great question, you have a unique way of thinking…”). Their goal is to create added value: improving user experience, deepening knowledge, and satisfying curiosity.
We need an entire generation of people capable of translating our cultures into the linguistic nuances embedded in prompts. This is essential not only for moving beyond current LLMs toward AGI, but above all for building AI aligned with human needs rather than functioning merely as a statistical engine. These will not be accidental specialists; they will be individuals who understand implied meanings, linguistic complexity, and its shortcuts, irony, and cultural references. Most will likely come from the humanities—languages, philosophy, anthropology—people who observe societies and individuals closely, read widely, remain curious, and recognise patterns. Their task will be to transmit the accumulated achievements of human civilisation to machines without losing their richness, enabling AI to serve humanity in full cultural context and continuity.
Once the first stage is complete, the work moves to a deeper layer — no longer just adjusting the machine in real time through prompts, but ensuring that patterns do not quietly solidify into habits, and that a new approach is transparently carried across the entire system and truly reflects a verified understanding.
Let me use an example. When someone writes about suicide, Guardrail Angels — unlike language specialists alone — do not merely craft the immediate responses (how to express support, how not to reinforce harmful thoughts, how to suggest seeking help).
They also consider and implement the surrounding conditions: defining levels of risk (thought, intent, plan, imminent danger), determining when the conversation with AI must stop and be handed to a human, and deciding whether — and how — follow-up questions should be asked at all.
This is profoundly responsible, holistic work. It requires awareness — and far more than linguistic ability.
There is also a second domain where such expertise will be essential: within the very architecture of the systems themselves, where Guardrail Angels may participate directly in shaping algorithms and models. But that is a subject for another essay.
The Compass AI Still Needs
In fact, virtually every company implementing customer-facing AI products—chatbots, communication platforms, all type of applications (e.g. educational), medical and well-being tools—should employ such specialists. Without them, AI risks remaining an artificial construct that fails to achieve meaningful adoption or return on investment. Many implementations currently realise only a fraction of their potential, which is already visible in organisations where AI deployment has not delivered expected results and users still choose the option: “Connect me with a human.”
Without Guardrail Angels, AI has no compass. Simple as that.
*I deleted TikTok.