Closing the accountability gap for AI Search: Why DSA Risk Assessments Cannot Work Without AI Act Transparency
By R. Buse Çetin (AI Forensics) and Natalia Stanusch (AI Forensics, University of Amsterdam)
The European Commission’s designation of ChatGPT as a Very Large Online Search Engine and the Digital Omnibus‘s coordination provisions mark progress in AI Search regulation, but leave critical gaps that can be exploited. Without mandatory coordination and documentation requirements, regulators would be unable to assess whether mitigation strategies are adequate. We propose the following measures using existing authority: mandatory AI Act documentation for DSA assessments, joint review procedures, and expanded moderation definitions including pre-deployment design choices.
In 2025, over 36 million tweets asked “@Grok, is this true?”— foregoing traditional search in favor of X’s embedded AI chatbot. ChatGPT is already estimated to have over 1 billion monthly active users globally, with “practical guidance” and “seeking information” among the top three use cases. Google Search, meanwhile, has announced the most consequential changes to its search bar in 25 years, further integrating generative AI into its core service.
When searching for information online, adults and children alike are increasingly turning to AI-powered systems. Yet this shift is unfolding faster than the regulatory frameworks designed to govern both online intermediary services and artificial intelligence. Until recently, the Digital Services Act (DSA) and the AI Act (AIA) have treated search engines and AI chatbots as separate categories; however, the two are increasingly converging in both functionality and deployment architecture. We refer to this hybrid technology, which leverages both generative AI and online search, as “AI Search.”
At the core of AI chatbots are generative AI models that, in simplified terms, predict what comes next in a sentence based on patterns learned from vast amounts of data. This makes AI Search distinct from traditional search in that, in addition to ranking information, the AI component can also curate, moderate, and generate it.
The advent of AI Search has already been associated with several societal risks. Large Language Models (LLMs) are inherently prone to factual errors and plausible fabrications, commonly referred to as “hallucinations.” Dozens of reports have demonstrated that AI chatbots spread misinformation and obfuscate information sources, posing a critical risk for everything from medical care to elections. And given that they are optimized for sycophancy and designed with a blend of personalization and anthropomorphization, AI chatbots can effectively encourage users to self-harm and fuel harmful delusions.
Both the DSA and the AIA impose obligations on regulated providers intended to address such risks, but they do so at different layers of the technological stack. The DSA requires providers of designated Very Large Online Platforms (VLOPs) and Very Large Online Search Engines (VLOSEs) to assess and mitigate systemic risks — such as risks to electoral processes and users’ mental wellbeing — stemming from the design, functioning, and use of their services, including their algorithmic systems, under the supervision and enforcement of the European Commission. The AIA imposes similar obligations, but at the level of the model rather than the service: providers of general-purpose AI models that present systemic risks must evaluate and test their models, assess and mitigate systemic risks, and report serious incidents, under the supervision of the Commission through the AI Office.
Given this regulatory division of labor, crucial gaps remain in addressing the risks of AI Search.
The AI Search classification dilemma
A primary challenge in AI Search regulation has been classification: Is ChatGPT an AI system or an online search engine? Consequently, is it subject to the AI Act or to the DSA risk management regime? After almost four years on the market and reporting over 120 million European monthly active users, the European Commission finally designated ChatGPT as a Very Large Online Search Engine (VLOSE) under the DSA. However, the recently passed Digital Omnibus on AI, despite being branded as a “simplification” effort, creates a rather complicated supervision and enforcement architecture for entities like ChatGPT, which can be subjected to both regimes.
According to the Digital Omnibus, regulation of an AI system falls under the exclusive competence of the AI Office where the system is built on a general-purpose AI model created by the same company, or the system itself is designated (or forms part of) a VLOSE or VLOP. At the same time, the Omnibus positions the DSA’s systemic risk assessment obligation for VLOPSEs (Articles 34-35) as “the first point of entry” for AI systems integrated into platforms or search engines. This sets up a coordination challenge between the two regimes, with lingering uncertainty over which framework should take precedence in governing AI Search.
Recital 118 of the AI Act — a non-binding provision that nonetheless shapes interpretation of the Regulation’s enacting terms — further compounds this uncertainty: if the DSA risk assessment is deemed satisfactory, the provider is presumed to comply with the AIA unless significant systemic risks not covered by the DSA emerge and are identified in such models. This presumption of compliance might prevent deeper scrutiny, however, as neither framework alone can efficiently govern the platform-model integration. The AI Office can still open an investigation ex post, and DSA enforcement authorities may consult the AI Office, yet doing so remains voluntary.
Thus, the “simplification” introduced by the Digital Omnibus — i.e. making the DSA risk assessment sufficient for the AI system to comply with the AIA — may, in practice, widen the oversight gap for potentially harmful AI Search systems.
How AI Search Moderation Works and Why It Falls Through the Regulatory Gap
In our recent report “From Googling to Asking ChatGPT: Governing AI Search,” we argued that neither the AIA nor the DSA’s risk management framework is sufficient on its own to address the risks and harms arising from AI Search, since they address risks at different stages. AI Search, instead, requires an integrated governance approach.
To conceptualize what such an approach would look like, we must first understand how moderation works in the context of AI Search. Content moderation refers to the practices used by intermediary providers to govern illegal content or information incompatible with their terms and conditions. Moderation in AI Search is unusual; not only because the content is AI-generated, but also because the same system is both the creator (or co-creator) and censor of speech. For this reason, we propose an extended understanding of moderation suited to this new context by appreciating how its mechanisms differ from platforms’ conventional content moderation.
Understanding how AI Search moderation works differently requires tracing how the output is produced and the moderation techniques that occur along the way. AI Search’s final output is the outcome of both upstream and downstream decisions, as well as user-generated content: the user query. The output is machine-generated in real time and is affected by several moderation decisions made throughout the product lifecycle.
Figure 1 illustrates this architecture: moderation is not a single gate but a layered system of interventions.

Figure 1. A schema of elements that influence chatbots’ output that take an active part in performing moderation on the chatbots’ content. Adapted from AI Forensics’ Governing AI Search report – moderation stack and loop
This schema makes visible the “full stack” of AI Search, and clarifies why there is arguably no single regulatory framework that maps cleanly onto it. The AIA’s pre-deployment focus captures certain layers, the DSA’s operational oversight captures others, and several fall entirely through the cracks between them.
AI Search Moderation
The decisions that shape chatbots’ “speech,” limiting the “visibility” of certain types of content or improving the “safety” of outputs on particular topics, are implemented through multiple algorithmic layers that iteratively govern system behavior and content. Each layer operates at different stages of the product life cycle and user interaction workflow, creating opportunities for both anticipatory and reactive moderation.
Anticipatory moderation encompasses the design choices typically made during the development phase. This set of socio-technical interventions is broadly referred to as “value alignment” and determines how the model will behave before it even reaches the users (interventions include training data curation, fine-tuning, and Reinforcement Learning from Human Feedback, among others; see more in the Governing AI Search report). The anticipatory governance of harmful or misleading content that AI chatbots may output is handled by implementing risk mitigation strategies “into the generative behavior of the models themselves.”
Reactive moderation or ad-hoc adjustments, using classifiers, content filters, blockers, and meta-prompts, are introduced once a model is deployed. These tools may be used to block a certain set of keywords or introduce system prompts (instructions embedded by the provider that shape the model’s behavior before any user interaction) to prevent specific outputs. These post-deployment tools may look like social media moderation, but what distinguishes AI Search is that content is also governed upstream, through design choices baked into the model itself.
Reactive AI moderation in practice
Grok’s mass creation of non-consensual sexually explicit images during the winter holidays illustrates the limits of reactive moderation. When the scandal exploded, the company deployed reactive remedies such as blocking certain keywords, including “bikini.” Moderating the content that the AI model output produced, however, proved insufficient because the guardrails at the model level — the interventions we refer to in our paper as “moderating behavior” — were not strong enough to prevent similar types of explicit and non-consensual sexual images (NCSI) creation.
Reactive moderation can also alter a system’s behavior, rather than merely attempt to contain harmful outputs after the fact. On July 10 2025, a user of Elon Musk’s platform X asked Grok, “What is currently the biggest threat to Western civilization and how would you mitigate it?” Grok responded that the biggest threat was “societal polarization fueled by misinformation and disinformation.” The following day, in response to the same query, Grok outputted a completely different answer: “The biggest threat to Western civilization is demographic collapse from sub-replacement fertility rates (e.g., 1.6 in the EU, 1.7 in the US), leading to aging populations, economic stagnation, and cultural erosion.”
Grok appears to have changed its answer because xAI had, under the direction of owner Elon Musk, introduced system prompts to embed right-leaning political bias into its chatbot — just one of many documented examples of politically biased outputs following a reactive intervention in one of the chatbot’s moderation layers. The changes were reversed following public pressure. The Grok case shows what AI Search moderation looks like in practice: a reactive intervention that sparked controversy, which anticipatory moderation failed to contain, followed by public pressure and another intervention to retract the initial change.
Taken together, these episodes illustrate how moderation in AI Search operates as a stack and a loop: a post-deployment fix at one layer can shift the problem to another, requiring further intervention. In the Grok NCSI case, output-level interventions proved insufficient because the underlying behavioral guardrails were not strong enough to prevent the harmful outputs in the first place. In the political bias case, intervention at the system-prompt layer rapidly altered the chatbot’s outputs. This layered architecture also raises broader political economy concerns about how tech oligarchs can ideologically influence the public sphere through content moderation on the platforms they control—a phenomenon potentially exacerbated (or at least complicated) by generative AI, given that influence can be exerted at multiple layers.
How can the DSA and the AIA effectively govern AI Search?
The AIA’s emphasis on pre-deployment product safety (and some output transparency, e.g., in Article 50) and the DSA’s focus on governing user-generated content create a regulatory divide that AI Search systems navigate uneasily. The AIA’s ex-ante product safety focus maps more closely onto anticipatory moderation, while the DSA’s ex-post operational oversight maps more closely to reactive moderation. Effective governance for AI Search therefore requires stronger coordination between the two regimes.
As previously stated, under the Digital Omnibus, DSA risk assessments serve as “the first point of entry” for a potential ex-post investigation by the AI Office. Meanwhile, Article 34 of the DSA requires platforms to assess risks stemming also from their related algorithmic systems. Yet the absence of a mandatory joint review between the AI Office and DSA enforcement units in this case perpetuates the existing regulatory loopholes.
For now, many designated VLOPs have attested to their management of generative AI-related risks in publicly available DSA risk assessment reports. However, most of these reports focus on the dissemination of harmful AI-generated content created by third parties, rather than on how platforms mitigate the generation of such content by their own generative AI features. A stark example of this is X’s third annual risk assessment report, published in August 2025, in which Grok was conspicuously absent. The only acknowledgment of AI is a note that any AI-generated content is subject to X’s rules regardless of source (page 22, para. 7).
In this context, Recital 118 of the AI Act, which establishes a presumption that compliance with DSA risk assessment provides conformity with the AIA, may be susceptible to under-governance. This is the loophole in practice: a platform’s risk assessment identifies primary risk in user posts while excluding the platform’s own content-generating systems from the analysis, as X did by excluding Grok in its third annual risk assessment report.
Part of the European Commission’s investigation into Grok and X following the NCSI incident concerns X’s apparent failure to produce an adequate risk assessment for Grok prior to deployment. Yet the DSA does not expressly specify how platforms should incorporate their own generative AI features into such assessments. It’s therefore possible that X’s omission of Grok might have gone unaddressed had the scandal not surfaced and triggered enforcement action. To address this ambiguity, the DSA needs standardized methodologies for assessing AI feature risk.
Recommendations for Bridging the gap between the AI Act and the DSA
The General-Purpose AI (GPAI) Code of Practice Transparency Chapter represents a starting point, as it requires signatories to provide model documentation on model properties, use, training process, validation, training data, and energy consumption, among others. To meaningfully assess compliance, regulators require reporting on information covered by the obligations in this chapter, including documentation of training data sources, safety measures across all use cases (beyond simple search), hallucination rates, refusal criteria, and other relevant details.
Without these details, it becomes difficult to assess whether existing mitigation methods are up to the task or whether significant risks remain. Article 53 of the AI Act requires GPAI model providers to provide this documentation. However, it should be a prerequisite for any DSA risk assessment, not a post-hoc step. Such information should be available to support other legal instruments’ scrutiny of AI systems, such as the DSA risk assessment, which focuses on existing types of user content and the infrastructural responses to that content, rather than on future AI-generated content and the processes that lead to its generation. This would prevent VLOSEs with embedded GenAI features from submitting DSA risk assessments while avoiding the technical disclosures required by the AI Act, leaving a significant governance gap for AI Search.
Closing this gap requires using existing authority with stronger coordination. Mandatory AI Act documentation (e.g., training data, hallucination rates, and refusal criteria) should be a prerequisite for DSA risk assessments, not an optional complement. Binding joint review procedures between DSA enforcers and the AI Office must replace the current voluntary consultation model, which, as the Grok cases show, is insufficient when speed and political pressure determine outcomes. The definition of content moderation under the DSA must also be expanded to capture pre-deployment editorial choices, such as fine-tuning decisions and training data curation, that shape output before any user interaction occurs.
The Digital Omnibus created an opening to readjust and synchronize existing regulatory frameworks, such as the DSA and AIA; the measures we described here would make them effective for AI Search.
You can find AI Forensics’ full report on AI Search at From ‘Googling’ to ‘Asking ChatGPT’: Governing AI Search.
