Two Stories of DSA Data (Non)-Access from an NGO perspective

By Oliver Marsh

This piece relates firsthand experience by the NGO AlgorithmWatch with multiple data access requests under the DSA – an early Article 40.4 request to understand AI content, and a “mass data access request” under 40.12.  It reflects on the (sometimes surprising) burdens encountered, and their impact: turning attempts to do research into full-time fights about compliance.


Data Access seems to be a persistence game more than it is a research method.  This piece relates a series of stories evidencing that theme.

To start with one story: The German NGO AlgorithmWatch has been trying, for some years, to understand the impacts of Generative AI on the information environment.  We had conducted research on the percentage of AI answers about German elections which contained issues such as inaccurate information. But we didn’t know what these percentages were of – i.e. what were the actual, real numbers of prompts that people were making about elections?

So in February 2024 – two and a half years ago – we submitted, to Microsoft and Google, a request to understand how many people had prompted Copilot and Gemini with terms related to German elections. Although the DSA was in force in 2024, at this stage we were still awaiting a Delegated Act which would clarify how 40.4 would work in practice. As such, DSCs were not prepared to accept an application, and we had to go directly to platforms. Microsoft actually provided us with data, though under heavy constraints – you can see more in Section 1 of this report. But even with limitations this was still than with Google, who refused (and much better than OpenAI, who just tried to sell us their enterprise product).

Though admittedly over-optimistic in its timing, the attempt was valuable in preparing us for the actual 40.4 experience. We learned that these requests require a lot of up-front preparation, as the chance to iterate after receiving early rounds of data is limited – a significant issue when trying to collect data from unpredictable and black-boxed systems. But most importantly, it also made clear that regulatory backing is needed, as otherwise researchers are negotiating backed solely with the goodwill of companies.  Fortunately that regulatory backing can now be provided under the DSA – in principle.

Data Access in the DSA

Access to data has been one of the key promises of the DSA. At its most basic, the DSA could help halt the trend of companies blocking data access – whether by ending free API access or tools like Meta’s Crowdtangle, or by legal threats against unofficial access. At its most exciting, the DSA offers access to internal company data, including valuable information on risk assessment metrics, content moderation, monetization, and the like.

briefly summarize the DSA Data Access rules: Article 40.12 allows researchers to request publicly available data directly from platforms (we made a one-page how-to here). Article 40.4 allows researchers to apply for non-public data via national regulators (Digital Services Coordinators, or DSCs) – see here for guidance from the Irish DSC. As of July 2026, 9 months after Article 40.4 came fully into effect, there had been 77 applications but no known successes. In both cases there are limitations to ensure the research is public interest (not commercial) and directed at addressing “systemic risks to the European Union”, and the applicants will be sufficiently responsible with the data (especially any personally identifiable data).

The DSA is still, arguably, the current best example of a data access regime for online services. The US failed to pass the Platform Accountability and Transparency Act; the UK is still working out how to apply its new Data Use and Access Act. But that is cold comfort for many researchers who are struggling to use the DSA to get the data we need. This is due to process issues and burdens which (still) suck a great deal of time and energy away from researchers and organisations already severely struggling for resources. This can be seen in our attempt to access quantitative data under Article 40.4 – a story I will return to shortly, and one that both echoes and diverges from the experience already described by Catalina Goanta and Anda Iamnitchi for this publication. Similar issues can also be seen in our experience of 40.12, including a collaborative attempt at providing ongoing access to the most-viewed posts across EU Member States on multiple VLOPs to a collection of organizations, which we call a “mass data access request”.

There is already plentiful writing about Article 40. There have been analyses of what data could be requested; what impact it has on public interest scraping (the Commission’s decision against X suggests there may be protections for some scraping); how companies impose limits via various terms and conditions; and tests showing technical issues with data provided.  This piece focuses more on the practical details of applications, the (sometimes surprising) issues these have raised, and the impact these have: turning attempts to do research into full-time fights about compliance.

Back to 40.4 – Google AI Overviews

In September 2025 AlgorithmWatch sent a complaint to the German DSC as part of an alliance of NGOs, associations, and media organizations. We argued, based on a review of emerging studies, that Google AI Overviews are a systemic risk to media pluralism – and that Google had not conducted an assessment of this risk, as required under the DSA. This complaint was forwarded to the European Commission (under the DSA, the main regulator of very large search engines like Google), who asked us to supply more data on the issue – in particular, around impacts on traffic to websites.

Attentive readers may spot that the Commission, as a regulator, could probably demand this information from Google itself. However, pointing out this issue did not change the matter. This continues in an unfortunate tradition of – we would argue – the Commission making substantial requests of civil society and researchers to support the DSA, without providing much information in return – though a new strategy and growing CSO engagement within DG-CNECT may hopefully lead to positive steps.

As such, when AI Overviews have been introduced into search results. We also asked for a copy of Google’s DSA risk assessments around AI Overviews. Although there is a small section on AI Overviews in their published summary (see page 66), it does not address risks to media pluralism. We wanted to ascertain whether there was any other assessments which did address this – as we would expect to have been carried out under the DSA. Importantly, we were clear that our request neither needed nor asked for personally identifiable information, thereby making the assessment of whether we had appropriate data protection easier.

As a German organization, we initially submitted to the German DSC; if we were completely ineligible, we felt this would be an easier way to find out. In December we received an initial positive assessment from them, confirming that the Irish DSC for detailed assessment of the actual request.  80 working days after submission – the deadline set by the Article 40.4 Delegated Act – we were rejected by Ireland.

But the grounds for rejection (or “revise and resubmit”, as we saw it) we received from the Irish DSC were reasonably detailed, bespoke, and surprisingly minor. They boiled down to two points. Firstly, some more specificity on our data storage and cybersecurity. Secondly, while we had showed the type of data requested was necessary for our research, we hadn’t explained why the amount of data requested was proportionate to our needs (with a balancing test against the interests of other actors involved). We have resubmitted, but because the Irish DSC currently assesses all applications in the order in which they’re received, we must wait another 80 working days.

40.4 take-away #1: The time it takes

The main take-away concerns how long it has taken. This stems partly from the 80 working day time limit for regulators.  However, shortening this in general would come with risks. The draft Delegated Act for 40.4 had a shorter time limit of 21 days. In feedback on the draft many respondents (including us) said this may be too short. DSCs can receive a wide range of potentially complex applications. For instance, the research topic may necessitate use of potentially sensitive data; there may be legitimate concerns about revealing too much about a company’s internal safeguards against malicious actors; a research team may have complex funding arrangements: and so on. It is good if such complexities are not insurmountable barriers to data access – but also that there is enough time for proper assessment.

However, in future, 80 days should be a time limit and not the expected period for everyone.  We made a relatively simple request for purely quantitative data – no GDPR implications at all. We still had to wait 80 working days for a decision from the regulator, and then were rejected.  At present, this is partly due to the Irish DSC assessing all applications in the order they receive them; so even though our application should be relatively straightforward, it has to wait for more complex ones. This makes sense while setting up a new process. But in future we believe it will be essential to triage simple and complex requests. Complex requests should of course expect longer vetting periods, but this cannot be a default expectation for any request under 40.4.

In addition to the 80 working days, we were also advised to delay the application on a couple of occasions to wait for new rounds of data access Guidance to be published (Guidance can be helpful, but working with the DSA has often involved waiting for documentation). Further – unexpectedly substantial – delays also came from numerous malfunctions in the Data Access Portal, which certainly needs more user testing. Finally, there was also of course the time needed to discuss and implement changes to the application.  By the time we receive a new answer from the Irish DSC, it will have been almost a year since first submission – and that’s before any negotiating with Google begins. Some of these tasks are hard to reduce – reworking applications will always require some effort, for example. But it is important that regulators are monitoring and addressing frictions, as they can add up substantially.  Otherwise 40.4 requests will forever struggle to address risks and functionalities in a timely manner, and will always be developed for situations that were issues months or years in the past.

40.4 take-away #2: There is a place for NGOs in 40.4

Our initial approval by the German DSC was somewhat of a surprise, and indeed a relief. As an NGO rather than a university or similar research institute, we had always been told that we were in principle eligible.  But we had doubts this would be followed in practice, and that NGOs without the data protection teams, processes, and infrastructure (e.g. clean rooms) of universities would struggle to get accepted.

We can compare our experience with that of a request made by academic researchers across two Dutch universities, which was also rejected – as Catalina Goanta and Anda Iamnitchi have usefully described for the DSA Observatory. Their request spanned a wider range of data from TikTok, to understand systemic risks arising from monetization of political content. Despite being based at a university, and proposing data security plans aligned with familiar academic standards, they were also rejected by the Irish DSC – on a wider range of grounds than we were, reflecting the wider breadth of their request.

This experience illustrates how many factors are at play in the outcomes of data access requests can be; and, as the authors note, the need for requesters to understand the legal and regulatory nature of the process. However the comparison between our experiences also illustrates a positive feature of 40.4, which had been promised on paper but may also be working in practice – that requests are assessed holistically, such that simpler requests may mean less strenuous expectations on data protection. Or, to put it another way: even if NGOs may struggle to meet the security standards of universities, they may still be eligible for 40.4 requests that don’t need those standards.

As more information on other requests emerges, and analysis from organizations like the DSA40 Data Access Collaboratory, we will be able to corroborate this further. But in the meantime, although 40.4 assessments involve work, we argue it is worthwhile for NGOs and other non-research-intensive organizations to lean towards applying. There are, of course, risks to overburdening regulators.  Keeping requests simple, or multiple organizations collaborating on requests, may help.  But as processes develop – and maybe even as data emerges – they need to be built for usable by a range of organizations, including NGOs.

A Mass Data Access Request under Article 40.12

Article 40.12 allows access to publicly available data, and has been in force – and therefore tested – for longer than 40.4. It has become a relatively “standardized” form of data access. In most cases, a VLOP or VLOSE has an online form (sometimes hard to find, though the DSA Collaboratory compiles a list). This gives – if successful – access generally via either a dedicated interface (such as Meta Content Library), or via an official API, or sometimes explicit permission to scrape.

However, Article 40.12 – despite its nickname as the “Crowdtangle Provision” – is not just about platforms providing data via their own official tools. The aim is to allow qualified researchers to access publicly available data through whatever means they prefer. For instance, after much hinting, the European Commission revealed in their fine against X.com their view that, if researchers can demonstrate they are meeting the requirements of 40.12, then VLOP/SEs should not block any scraping they conduct. However, while scraping may give flexibility, it can also be challenging for organizations with less technical capabilities. Other 40.12 possibilities could include, for example, VLOP/SEs providing dedicated datasets.

The wording of 40.12 also allows for more creative approaches to accessing publicly accessible data. In April 2024 AlgorithmWatch, the Mozilla Foundation, and the DSA40 Data Access Collaboratory jointly organized a “data access hackathon” to explore these options. While multiple options were mooted, we decided to first try an apparently simple idea from technologist Louis Barclay – a request for “highly viral content”.

Making the Mass Request

We brought together roughly 20 research organizations into a joint request to six VLOPs (Facebook, Instagram, YouTube, TikTok, LinkedIn, X.com), asking them to share the top 1000 most-viewed posts in each Member State every six hours. As such highly viral posts reveal a lot about narratives trending on a given platform – with accompanying political and civic importance, and also risks of disinformation – such data would play a valuable role for research, monitoring, and fact-checking organizations. Also, we argued, being viral means such content should clearly be considered publicly accessible.

We took a collaborative approach to substantially reduce burdens on both the organizations and the platforms. We wrote one justification as to why the research itself met the requirements of 40.12 (such as how it was related to systemic risks), and appended documentation from each individual requesting organization to show they met the independence, data protection, etc. requirements. This documentation was sent once, by email, to relevant email addresses at each VLOP. This approach saved each organization needing to individually submit – and each VLOP receiving in their assessment inboxes – multiple near-identical applications. By asking the platforms to provide the data directly, we also avoided the need for multiple identical (and repeated) requests to platforms’ APIs or dashboards from all the different organizations. It has involved a lot of up-front central coordination work from the three leading organizations, but if it works, it will be a far more efficient approach than following the “expected” route.

Results: More Waiting

We have now been in email back-and-forths with the platforms since October last year. None have so far acceded to the original request, usually with similar reasons. Some of the platforms dispute that the data is publicly available, despite being highly viral posts. In some cases this hinges on technicalities. For instance, if view counts are not public, we cannot request data which is sorted using these non-public view counts – even if we never actually receive the view counts. Other platforms have stated we should apply using the “official” route, or use the Article 40.4 route, despite the huge additional burdens this would present (also, some of the tools provided by “official” routes cannot actually perform the task we need). But we have been willing to negotiate on these points.  We are continuing these email exchanges – except with X.com, who bluntly told us they would not reply to further emails – but are not, at present, hopeful the 40.12 route will work as we had hoped.

We suspect the underlying concern from the companies is that this work could illustrate how highly viral posts spread polarization and inaccurate information – as Crowdtangle revealed, before it was shut down by Meta. But this undermines the essential purpose of 40.12: giving responsible researchers the chance to see if platforms do have problems they should be addressing. The basic data we are requesting – highly viewed posts – cannot reasonably be called “not public”.  In this case, the imbalance of power that regulation should solve is not working.

We have been running a petition to illustrate the interest in this data.  If the VLOPs nonetheless continue to refuse us, this then leaves us with the route of complaining to the European Commission. But then we end up in the familiar DSA situation – unclear if, how, or when the complaint will lead to results. And all this to try and access a fairly simple, but important and clearly highly public dataset.

Conclusion: When do we go from teething problems, to regulation with teeth?

These experiences continue in a long line of burdens faced by researchers trying to hold large online service providers accountable – including into the fourth year of the DSA. The aim of relaying these experiences was to provide examples of what these burdens can look like in practice, how often we encountered them, and the impact they have in aggregate. Burdens have arisen across the different forms of data access. For 40.4, we have kept the request simple, aware we are using a new and in-development process. For 40.12 we are trying something more innovative, to go beyond the limitations of the provided tools; but the aim is still very aligned with the aims of 40.12 – accessing very public data – and carefully designed to reduce complications for all concerned. In both cases, coordination and admin work has far exceeded time spent on research planning or execution.

Implementation of regulation always has teething problems, and always requires testing and learning. The DSCs in particular were helpful in explaining their constraints and approaches, and 40.4 is new and complex – but also important and needs to work. But for 40.12 we have now had over three years of teething problems, and are still facing fairly basic issues.

There is a risk of repeating a similar situation from across in the DSA, seen for example in systemic risk research.  There is awareness of the great potential for support from researchers and civil society, a realization that the current situation isn’t working, hopes that there will be improvements over time, but no clear roadmap or transparency from authorities on how this will happen. What external researchers need are (i) clear, open, and concrete plans from regulators on how they will address problems with DSA data access including (ii) making timely warnings, and enforcement if needed, against platforms who are resisting reasonable data access requests.

We have recently seen a flurry of outputs emerging from internal regulatory activity. Let us hope this energy can turn to bringing external actors in more effectively. Data access is necessary, and the DSA approach can still be a model.  But it must be a model that works in practice, or improves when it does not.

Thanks to LK Seiling of the DSA40 Collaboratory for feedback on an earlier draft.