As digital bullying becomes more sophisticated, is AI moderation effective? Do we need better human oversight?
New Face of Bullying
Digital bullying no longer sounds like bullying. It’s evolved, wrapped in sarcasm, using sophisticated language, exclusionary phrasing, or even professional politeness. What once looked like open hostility now hides behind passive-aggressive civility and edgy intelligence. The harm is often imperceptible, but no less real.
But are automated moderation systems struggling to keep up?
AI tools are designed to look for specific words which can be a flawed methodology and the cause of ambiguity:
- Certain words when legitimately used can be flagged for moderation and even result in valid content being removed completely (sexual harassment, sexism, women’s health issues and mention of female body parts etc.)
- AI allows the detection of explicit slurs or insults, but it misses the subtler forms of harm embedded in tone, power dynamics, and context. A comment might look respectful on the surface yet be deeply cutting or exclusionary in intent.
- AI can also misread frustration or irony as aggression, creating false positives that erode trust in moderation.
This is where human understanding becomes essential and even then only when people are trained. Research from the Harvard Kennedy School of Misinformation examines the challenges of identifying intent in various types of online abuse, for example, hate speech and cyber-bullying.
Algorithmic Design
Research also shows that social media algorithms can deliberately, or at least functionally, feed content with opposing ideological views into users’ feeds, not primarily to inform but to rage bait. By showing liberal content to right-wing accounts (and vice versa, such as feminist or progressive messages to conservative users), the platform amplifies conflict, triggering strong emotional reactions including outrage and anger.
This engagement loop feeds the algorithm: more conflict = more activity = more time spent on the platform, which fuels ad revenue and user retention.
Given that certain platforms intentionally create conflict shouldn’t they have effective moderation channels in place when some people cross the line?
My Case Study: When AI Misses the Point
Recently, I was stopped by LinkedIn for calling a well known UK politician a “grifter” in a comment. This is a factually accurate description, consistent with the dictionary definition, given that said politician had 50% of his salary deducted following charges of misuse of EU funds.
Meanwhile, another user re-shared one of my posts on algorithmic bias with this commentary:
“This nonsense could have been prevented by paying attention in science lessons aged 12…
Why am I not surprised an HR person says things like this?
Dorothy Dalton, how can you do diversity if you don’t understand the basics?”
He then continued to troll the thread with similar remarks, several of which I deleted.
When I went to report this behaviour, my only option on LinkedIn was “harassment”, defined thus
“Harassment: Attacks or intimidation toward others with abusive language or deliberately and repeatedly disrupting conversations, including revealing others’ personal or sensitive information. This includes unwanted romantic advances, sexual remarks and or requests for sexual favors.”
Definitions of Bullying
The ILO’s definition of bullying is:
Convention No. 190 (C190) states that bullying is characterized by a single occurrence or repeated instances of unacceptable conduct that can create an intimidating, hostile, humiliating, degrading, or offensive environment. It is not defined by actual harm but by the potential to cause harm.
To check myself, I ran my troller’s comment through ChatGPT for tone analysis. The system described it as:
“Critical and mocking, mixing sarcasm with frustration and condescension… the combination of mockery, superiority, and accusation fits the pattern of verbal or relational bullying, where language is used to undermine confidence or social standing.”
So it would seem that some AI can see nuances, but LinkedIn’s systems don’t. The platform also has no reporting category for this kind of behaviour. And it turned out that this individual had targeted several other women in similar ways.
It’s a small example, but it reflects a wider issue: many women hesitate to engage publicly online because the line between critique and cruelty has become so easy to blur, and so easy to ignore. The ineffective treatemtn of bullying by AI where women’s voices are silenced.

Gender and Power Dynamics
Online bullying doesn’t happen in a vacuum; it mirrors offline hierarchies and biases. What’s often dismissed as “robust debate” or “professional criticism” can actually be gendered or targeted behaviour designed to erode confidence and credibility.
1. Gendered Bullying Disguised as Professionalism
Women, especially those in visible or leadership roles, often experience hostility disguised as reasoned debate. Phrases such as “you don’t understand the basics,” “maybe stick to your area,” or “you’re too emotional” are presented as logical disagreement, but the subtext is about control or dismissal.
Because these comments are framed politely, both algorithms and human reviewers frequently overlook them. This allows aggressors to maintain plausible deniability while systematically undermining others.
Worth a read: Algorithmic Bias. The New Glass Ceiling
2. Tone Policing
It’s rarely one remark that silences women; it’s the steady drip of belittling, sarcasm, or “just joking” comments that add up over time. Tone policing (“don’t take it personally,” “you’re overreacting”) shifts responsibility from the aggressor to the target, making women question whether they are too sensitive rather than whether they are being mistreated. Standard DARVO tactics.
This erosion of psychological safety drives self-censorship and disengagement. The cost isn’t just personal, it’s societal. When certain voices withdraw, collective intelligence shrinks.
Consider this: Lead with Safety: Building Inclusive & Psychologically Safe Teams
3. Failed moderation
When moderation fails, marginalised groups pay the price. Algorithmic moderation systems already carry bias from their training data. When these systems fail to detect relational or coded bullying, the harm isn’t evenly distributed. Women, LGBTQ+ users, and professionals of colour bear the brunt.
Research by the Center for Countering Digital Hate found that automated moderation fails to detect over 70% of abusive content directed at women, particularly when it’s phrased in coded or sarcastic ways.
In other words, the smarter the bully, the safer they are. Those who can weaponise politeness or intellectual superiority exploit these blind spots, while those calling out bias risk being flagged themselves.
When moderation systems fail to recognise these nuances, platforms don’t just overlook bullying; they reinforce existing power imbalances.
Other Platforms
Different platforms handle this challenge in vastly different ways, often with equally mixed results.
Reddit relies heavily on volunteer moderators who interpret community norms and apply context. This human element allows nuance but creates inconsistency across subreddits.
TikTok uses large-scale AI moderation for speed but has faced criticism for inconsistent enforcement and cultural bias, particularly around gender and ethnicity.
Twitter (Never X) has rolled back much of its Trust & Safety infrastructure, removing teams focused on harassment and misinformation. The result? A rise in abusive content and coordinated trolling, particularly targeting women and journalists. The platform is now a cesspit, and once a Twitter addict, I am no longer active.
These examples show that automation alone isn’t enough, but neither is untrained or inconsistent human oversight. The challenge is balance: scaleable systems informed by empathy, context, and expertise.
Why We Need (Trained) Human Oversight
This is about more than missed keywords; it’s about a failure to determine intent.
Humans are uniquely equipped to detect patterns and relational cues: the sarcasm, the repetition, the small shifts in tone that indicate someone is being targeted over time. Algorithms can surface potential issues, and human reviewers can interpret context, nuance, and cultural dynamics, especially if they are trained to identify bias.
The goal isn’t to replace AI with humans, but to combine the best of both. AI provides scalability; humans provide empathy and judgement. Together, they form a “human-in-the-loop” model where automation filters content and trained experts make contextual decisions.
Ethical and Design Implications for Human Oversight

If we accept that AI alone can’t manage sophisticated bullying, then we must ask: what does responsible human oversight actually look like?
1. Training in Bias, Culture, and Gender Dynamics
Human moderators need more than procedural checklists; they need cultural and emotional intelligence.
Training should cover implicit bias, gendered language, and regional communication norms to help moderators accurately interpret intent. A phrase that sounds neutral in one context can be deeply patronising in another. Recognising these subtleties is what turns moderation from mechanical policing into digital empathy.
2. Accountability and Transparency
If we introduce human oversight, it must also be accountable. Who selects moderators? How are their decisions audited for bias? Can users appeal decisions? Without transparency, even human moderation risks perpetuating the same inequities it’s meant to solve. Clear reporting structures and open feedback channels can help build user trust.
3. AI and Humans as Co-Learners
The most effective systems treat humans and AI as mutual teachers. Human moderators can label complex examples — sarcasm, coded exclusion, condescending “helpfulness, ” so AI systems can learn to recognise them. Conversely, AI can detect emerging language patterns or coordinated behaviour that humans might miss. This creates a feedback loop of ethical learning where both technology and people improve together.
4. Trust Not Compliance
Ultimately, moderation should not be about simply removing harmful content; it should be about protecting community trust. That means shifting from a rule-enforcement mindset to one of care and accountability.
Trust & Safety isn’t a compliance function, it’s a culture function. And that’s where human oversight becomes irreplaceable.
5. Agreement around definitions
It’s important that platforms use recognised legally accepted definitions rather than create their own. LinkedIn’s definition of harassment is inadequate and offers no possibility to report bullying.
Moving Forward
As online communication grows more linguistically and socially complex, Trust & Safety systems must evolve, not just technologically, but humanly. We need moderation models that combine AI’s efficiency with human insight, and that recognise the lived realities of those most affected by digital hostility.
Because when subtle bullying goes unchecked, it doesn’t just harm individuals, it discourages people from particpating thus narrowing the space for genuine dialogue, inclusion, and innovation.





