Perspective on responsible use of AI for analysing large document datasets

Our working perspective on how AI helps across the research lifecycle in large-scale qualitative analysis, especially as the field shifts from episodic to continuous research.

Cover Image for Perspective on responsible use of AI for analysing large document datasets
Share this article:

This is a short note outlining my current thinking on the main uses of AI in knowledge work involving large bodies of text, prompted by conversations about AI use in the analysis of large document datasets: consultations and interview data in particular. It is organised around three complementary ways to think about how AI is used in public sector and commercial work:

  1. The lifecycle of the research or evaluation process
  2. The four types of cognitive tasks AI performs within that lifecycle
  3. The shift from episodic to continuous research that better tools make possible

1. AI across the research lifecycle

Research design and data collection

Research design is largely driven by expertise and client understanding. To facilitate human intuition, it is useful to analyse relevant existing datasets to devise a sampling strategy and interview guide or survey. With artificial intelligence it is far faster to take stock of previous datasets:

  • Spot themes that are specific to demographic groups, so you can sample sufficient respondents from each
  • Identify unexpected or interesting themes in prior data that warrant new questions

In the pre-AI era it was very difficult to take stock of previous datasets in any reasonable amount of time. That has now changed, and it changes what a good sample size and interview guide can be informed by before fieldwork even starts.

Data familiarisation

In qualitative research, the "framing" of a study influences how data is read, thought about, and reported. AI lets a researcher generate and assess alternative approaches rapidly by familiarising themselves with the data first. A quick AI pass can tell you whether it makes more sense to organise findings by the type of stakeholder, the stage of the process being studied, or the different outcomes observed.

Here, metadata is very useful: are there topics disproportionately discussed by some actors but not others? Tagging documents with stakeholder type, role, or process stage and cross-tabulating themes against that metadata turns "I have a hunch this differs by group" into something you can check before you commit to a framing.

Main data analysis: developing the narrative

Every qualitative research project needs a strong narrative. There can be separate themes, but the outcome is stronger if they can be tied together. The narrative itself is up to the human researcher to decide, but AI can surface themes quickly and identify how many respondents or source documents touch each one.

With advanced qualitative research tools, users can quickly recategorise the data from alternative perspectives, rapidly identifying the optimal way to structure equivocal material, rather than living with the first category structure that emerged. That is the practical version of what we call method-agnosticism in our design criteria for AI-assisted qualitative analysis tools: a tool should let you try on a different frame cheaply, not lock you into the first one. See how to synthesise research findings into a narrative for more on that step specifically.

Power quotes, robustness, and counter-evidence

Humans suffer from confirmation bias; we want to see the results we expect in the data. AI is a useful co-worker here precisely because it has no stake in the outcome. AI tools can assess the robustness of findings by systematically evaluating the presence of aligned evidence across all documents, and they can help surface counter-evidence, either in the form of contradictory statements or by identifying alternative explanations. We have written before about why rigour means being willing to return a negative result, and this is that discipline built into the workflow rather than left to individual researcher virtue.

Finally, in the later stages of analysis, AI can help identify the clearest quotes, often called "power quotes", that work well in reports and slides because they exemplify a broader pattern in the most succinct way possible. The precondition for trusting a power quote is that it is actually there in the source: see a citation is not an analysis for why a verified quote and a defensible finding are not automatically the same thing.

Research report

The final step of research involves composing a report. AI can speed up synthesising the data and writing the text around the quotes. This is where AI slop is a real problem. Keeping original quotes attached to every claim reduces hallucination risk and lets a human verify that the AI's synthesis is actually aligned with the data, rather than a plausible-sounding gloss on it. Our posts on hallucinations, context and the black box and two-way transparency go into the mechanics of why that traceability matters more than it might seem.

2. Four types of cognitive tasks

The lifecycle stages above map onto a smaller set of underlying cognitive tasks, and each one carries its own risk profile.

TaskWhat it doesKey risks
Bottom-up codingReads text at a low level to identify passages relating to a theme of interestFalse negatives, false positives, systematic bias against non-standard phrasing or under-represented groups
Top-down categorisationDigests large lists of quotes and divides them into broad, often hierarchical themesBland, self-evident output; over-accommodation of all content rather than focus; miscategorisation; distinctive themes hidden inside broad ones
SynthesisingBuilds structured syntheses with meaningful linkages between findingsOverinterpreting scarce evidence, spurious linkages, selective emphasis on self-evident aspects
Structured comparisonsCompares results across time or across groupsHallucination, overinterpretation of differences from individual observations

Bottom-up coding

AI can be used with a low-level focus on text to identify passages and paragraphs relating to themes of interest, letting the researcher capture all relevant passages across large bodies of material rather than the subset they had time to read closely. See our guide on how to code qualitative data for the inductive mechanics this task rests on.

Risks: false negatives (missing content that should have been captured), false positives (overinterpreting content to fit the category in question), and systematic bias, for example problems faced by a specific social group, or expressed in non-standard English, being undercounted.

Top-down categorisation

AI can digest relatively large lists of quotes and divide them into broad themes. Software like Skimle can process quotes in batches and further construct subdivisions of key themes into subthemes, building a hierarchical category scheme not unlike the open-to-focused progression in grounded theory. Ministries in Finland found that AI-generated thematic categories provided a useful basis for conversation, and even for collecting further feedback once the categories existed. A similar dynamic showed up when we analysed 500+ responses to the EU Digital Omnibus consultation: the category structure itself became something policymakers could react to, not just a summary they read once.

Risks: categorisation ends up bland, self-evident, and descriptive rather than something that supports a decision. Effective categorisation focuses on a subset of the content; AI typically tries to accommodate all of it in a parsimonious way instead. Categorisation can also be biased or flawed outright, with miscategorised material and distinctive themes hidden inside categories too broad to see them.

Synthesising

AI has long been used to summarise content, but it also become better at creating structured syntheses with meaningful linkages between data points. The AI's ability to synthesise empirical findings well depends on meaningful instructions, whether those come from an agentic system, a data pipeline, or a human expert. Even with the most advanced LLMs, still today if you just try to upload an unstructed dataset the AI produces syntheses that are not trustworthy. With Skimle we take steps to first analyse and structure the data before attempting synthesis.

Risks: without a dedicated workflow, syntheses tend to overinterpret scarce evidence, create spurious linkages between unrelated data points, and selectively emphasise the self-evident parts of the data rather than the parts that actually needed synthesis.

Structured comparisons

AI is particularly good at finding patterns within qualitative categories over time or across groups, for example, do the claims related to unclear targets differ between 2024 and 2025? When running these comparisons, we have found that the less you tell the model about the groups being compared, the better. If you compare quotes from men and women and simply tell the model they are group A and group B, you get meaningful results. Name the groups explicitly, and the model tends to hallucinate stereotypical differences instead of reporting what is actually in the data. It is the same finding behind criterion 7, adversarial rigour, in our design criteria post: S kimle runs segment comparisons as "group A" versus "group B" for exactly this reason.

Risks: hallucination, and overinterpretation of differences that rest on a handful of individual observations rather than a real pattern.

3. From episodic to continuous research

The high cognitive cost of familiarising oneself with data meant it was impractical to collect data on several themes in parallel and analyse them over time. With AI, it is now possible to automatically code data against an established category structure as it arrives. That lets AI-empowered processes identify changes in trends and new emerging topics as they happen, and within-category comparisons over time can show how the detail of a specific finding shifts, not just whether the headline theme is still present.

In private sector customer experience and market research, several companies are pushing to shift from episodic studies to continuous data collection and analysis, precisely to better understand weak signals and to evaluate the impact of interventions and development programmes as they unfold rather than after the fact. Our posts on always-on customer research and NPS verbatim analysis at scale cover what that looks like in practice, and combining insights across feedback channels covers the harder version of the problem, where the continuous data is not all coming from the same source.

The mechanistic nature of AI coding also means that qualitative analyses conducted over time are more uniform and comparable than analyses done by different humans at different points in time, each primed by the surrounding agenda and public discourse of that moment. Changes in AI models can create drift of their own, but that is a more tractable problem: it is easy to re-analyse historical data with the newest model to maintain comparability, in a way that re-running a study with a different human panel of coders never quite is.

So what: how to incorporate AI into your knowledge work?

Hopefully these three lenses are useful in developing your own processes and competences.

The lifecycle view offers a way to think about concrete process flow: it is worth discussing, stage by stage, the extent of the work involved and AI's ability to both speed things up and increase the quality of outputs by letting experts do things they never previously had time for.

The cognitive tasks view is perhaps the most useful for developing human skills, because these task types cut across the whole research process rather than belonging to one stage. Each type of AI use carries its own human oversight needs and risk management approach, and those are things your team will get better at judging over time rather than something you can fully specify in advance.

The episodic-to-continuous view raises the harder question of process redesign: as people get more competent with AI and the tools keep improving, should you rethink how the work itself is organised, and over what timespans? The aim, somewhat ironically, is to become both more responsive in the short term and more accountable in long-term strategic evaluation at the same time.

If you want the more general design argument behind why tools should be built this way in the first place, our posts on 9 design criteria for AI-assisted qualitative analysis tools and designing AI that augments qualitative researchers lay that out in more depth.

If your own work sits in public sector consultation analysis specifically, see stakeholder consultation analysis for policy teams and analysing public comments in government consultations for the more procedural version of what is described here, and how Skimle fits public sector and policy work more broadly.


Want to see this in practice on your own document set? Try Skimle for free and run a batch of documents through bottom-up coding and top-down categorisation to see where the lifecycle stages above map onto your own workflow.

Related reading: See 9 design criteria for AI-assisted qualitative analysis tools and designing AI that augments qualitative researchers for the underlying design argument, and primary research for consulting projects for how the lifecycle view applies outside the public sector.


About the author

Henri Schildt is a Professor of Strategy at Aalto University School of Business and co-founder of Skimle. He has published over a dozen peer-reviewed articles using qualitative methods, including work in Academy of Management Journal, Organisation Science, and Strategic Management Journal. His research focuses on organisational strategy, innovation, and qualitative methodology. Google Scholar profile

Dig deeper to your data with Skimle

Skimle collects, analyses and categorises interviews, survey responses, reports and other qualitative data automatically. Our modern qualitative analysis software combines a rigorous and transparent workflow with the speed of AI.

Upload text or audio, remove sensitive data with Skimle Anonymise, automatically create categories and sub-categories, explore the data across documents and export the data to seamlessly fit your workflow. Built by professionals for professionals, with full privacy and GDPR compliance.

Free trial · No credit card required · Full plans from €20/month