AI can support ethnographic research at every stage without replacing the fieldwork itself: mapping existing theoretical framings during the literature review, transcribing field notes and interviews, surfacing candidate themes across observation data, and testing alternative ways to interpret a pattern. Participant observation, rapport and in-the-moment interpretive judgement stay with the ethnographer throughout.
Ethnography is the qualitative method built on watching what people actually do rather than what they say they do, and that gap between stated and enacted behaviour is exactly what makes it resistant to shortcuts. You cannot prompt a language model to sit in a factory canteen for three weeks. What you can do, carefully, is use AI to lighten the parts of an ethnographic study that are mechanical rather than interpretive: the reading that precedes fieldwork, the transcription that follows it, and the first pass through hundreds of pages of field notes before the real interpretive work begins.
This guide walks through where AI actually helps across an ethnographic study, using Skimle as the working example, and where it should not be anywhere near the process. It draws on the same design criteria we set out in 9 design criteria for AI-assisted qualitative analysis tools and the predictability, transparency and control framework from designing AI that augments qualitative researchers, because ethnography's reliance on situated, contextual judgement makes those criteria more important here, not less.
What makes ethnography different from other qualitative methods?
Ethnography is observation-first. The researcher spends extended time in a setting, participating to varying degrees, and produces field notes as the primary data rather than a transcript of a single conversation. Barbara Kawulich's widely cited framework for participant observation (2005) sets out a spectrum from complete participant to complete observer, and the researcher's position on that spectrum shapes what they can even see, let alone record.
Clifford Geertz's foundational term for the output is "thick description": an account that captures not just what happened but the layers of meaning participants themselves attach to it (Geertz, The Interpretation of Cultures, 1973). A thin description records that someone winked. A thick description records that the wink was a private joke, aimed at a specific person, referencing a shared history the researcher had to spend weeks in the field to learn about. That layered, context-dependent meaning is precisely the part no model can supply, because it does not exist in the text of a field note. It exists in the researcher's accumulated understanding of the setting.
This is the reason ethnography deserves separate treatment from AI-assisted interview or survey analysis. An interview transcript is largely self-contained. A field note is not; its meaning depends on weeks of context that never made it onto the page. Ethnographic data is also rarely just notes and interviews: photographs, documents, and physical artefacts collected from the setting often sit alongside them as evidence. Any AI tool used in ethnographic work has to be positioned as support for the mechanical layers of the process, never as a substitute for the fieldwork or the interpretation built on it.
Where does AI actually fit across the 4 stages of an ethnographic study?
A multi-week ethnography can easily generate tens of thousands of words of field notes and transcripts before a single theme is coded, and each of the four stages below has a different relationship to AI assistance.
| Stage | What it involves | Where AI helps | What stays with the researcher |
|---|---|---|---|
| Literature review | Mapping existing ethnographic and theoretical work on the setting or population | Surfacing themes and framings across dozens of papers quickly | Choosing which framing fits your research question |
| Fieldwork capture | Recording field notes, informal conversations, and any interviews conducted alongside observation | Transcribing audio accurately and consistently | Deciding what to observe, building rapport, in-the-moment judgement |
| Theme discovery | First-pass coding of field notes and transcripts | Surfacing candidate codes and categories across a large corpus | Judging whether a code captures the meaning, not just the words |
| Pattern and reframing | Testing whether findings hold across sites, roles or time, and considering alternative theoretical lenses | Cross-tabulating by metadata, restructuring categories under a different frame | Deciding which framing is the right one to publish |
The rest of this guide goes through each stage in turn.
How can AI help with the ethnographic literature review?
Before fieldwork starts, most ethnographers already have a stack of related studies: prior ethnographies of similar settings, the theoretical traditions those studies drew on, and the framings other researchers arrived at. Reading through that stack to understand which theoretical lens fits your own question, and where the gaps are, is closer to a synthesis task than to fresh analysis.
This is the same mechanical bottleneck we cover in qualitative evidence synthesis and meta-ethnography: once a literature review has moved past 20 to 30 papers, close reading and cross-referencing by hand slows down considerably. George Noblit and R. Dwight Hare, who developed meta-ethnography as a formal method for synthesising qualitative studies, worked with only 2 to 6 studies at a time in their original 1988 book, precisely because hand translation between studies is so labour-intensive. A widely cited follow-up analysis puts the practical ceiling for manual meta-ethnography at around 40 studies "to allow sufficient familiarity" with the material (Toye et al., 2014, BMC Medical Research Methodology). That ceiling is a feasibility constraint on hand extraction, not a methodological rule, and AI-assisted extraction is exactly what loosens it.
In practice, this looks like uploading the papers or their extracted findings sections as documents in a project and running an inductive pass to surface candidate themes across the corpus, the same operation described in the evidence synthesis guide. For ethnography specifically, this is most useful for mapping the theoretical framings other researchers applied to comparable settings before you commit to one of your own. It is a literature-mapping exercise, not a replacement for reading the studies that matter most to your argument closely yourself.
If you work in academic research, how Skimle fits an academic workflow covers this kind of pre-fieldwork groundwork alongside the reporting and ethics side.
Can AI transcribe field notes and interviews accurately enough for ethnographic analysis?
Ethnographic data collection is not limited to audio. Field notes are often written or dictated after an observation session, informal conversations get scribbled down in a notebook, and any formal interviews conducted as part of the study produce recordings like any other qualitative interview. Where audio exists, transcription accuracy matters as much here as anywhere else in qualitative research.
OpenAI reports a word error rate of around 2-3% for its Whisper model on clean English audio, a benchmark most commercial transcription tools now compete against; our comparison of AI transcription tools for research covers how the leading options perform on accented speech, technical vocabulary and background noise, all of which are common in field recordings taken outside a controlled setting. For dictated field notes (a researcher speaking observations into a phone immediately after leaving the field), the same transcription pipeline applies.
Skimle transcribes uploaded audio directly into a project, so a field recording becomes an editable, searchable document rather than a separate file you have to reconcile with your notes later. See uploading voice files for supported formats. Once transcribed, the same review discipline applies as to any interview transcript: read it against the recording before coding, because transcription errors that misattribute speakers or drop a qualifying clause change meaning in exactly the passages an ethnographer is trying to capture precisely. Dictating notes immediately after leaving the field, rather than writing them up hours later, is standard fieldwork practice for a reason: recall of specific detail fades quickly once a researcher steps out of a setting, so getting words down fast matters more here than in most qualitative data collection.
Privacy deserves more weight here than in most qualitative work. Ethnographic field notes routinely describe a specific workplace, neighbourhood or community, sometimes down to physical layout and named roles, which makes participants identifiable even after their names are removed. As Asher Beckwitt, a medical and cultural anthropologist who advises teams on responsible AI practice, puts it in the context we cite in our design criteria post: a participant "may remain identifiable through position, location, diagnosis, relationships, or a combination of demographic details" even once a name is gone. Anonymising documents before analysis, and thinking through what combination of descriptors could re-identify a specific site or person, matters more in ethnographic write-ups than in almost any other qualitative genre. Our guide on IRB-compliant pseudonymisation covers the practical steps. What no tool can settle for you is whether your ethics approval and the consent participants gave actually cover processing field data through a third-party service; that determination belongs to the researcher and the ethics board, made before fieldwork begins rather than after.
If your study also includes structured interviews alongside observation, whether recorded live or conducted asynchronously, transcription is the same operation either way. Consider trying this on a small batch of your own field recordings with a free trial before committing a whole fieldwork corpus to any single tool.
How does AI find themes and patterns in ethnographic field data?
Once field notes and interview transcripts are assembled, the first coding pass is where AI assistance saves the most time, and where the design criteria we set out in designing AI for qualitative research matter most: predictability, so the same corpus does not produce a different theme structure on every run; transparency, so every candidate theme traces back to a specific passage in a specific field note; and control, so recoding is a click, not a five-minute rerun.
Ethnographic coding usually starts descriptive (what happened, who was involved, what was said) and moves toward more conceptual, interpretive categories, mirroring the open-to-focused progression described in grounded theory methodology. Skimle's predefined categories mode supports both ends of that progression: an inductive pass over your uploaded field notes and transcripts to suggest initial descriptive categories, or a deductive pass against a theoretical framework you write yourself if your study is testing a specific theory rather than building one from scratch. Our guide on how to code qualitative data covers the inductive, deductive and abductive distinctions in more depth.
Metadata is where ethnographic coding differs most from single-session interview analysis. A study typically spans multiple sites, roles, observation sessions and weeks of fieldwork, and a theme that holds on day one but disappears by week four is itself a finding. Tagging documents with metadata such as site, participant role, session type and date lets you cross-tabulate emerging themes against those variables, and the timeline view shows whether a category grew, shrank or stayed flat across the fieldwork period. Watching a theme's prevalence shift as you spend more time in a setting is close to what ethnographers already do informally by re-reading their own notes; the difference is doing it systematically across the full corpus rather than from memory.
As with any AI-assisted coding, the categories view is where the judgement work happens after the first pass: renaming a code that does not capture what participants meant, merging two categories that turned out to describe the same underlying practice, splitting one that was doing too much work. Our two-way transparency post covers why being able to trace a theme back to source, and a source document forward to what was coded from it, is the check that separates a defensible thematic account from an assertion.
Can AI help you discover patterns and explore alternative ways to frame your findings?
This is where ethnographic work diverges most from a single, fixed coding scheme. A good ethnography usually tries on more than one theoretical lens before settling on the account it publishes: is this workplace practice best explained as resistance, as sense-making, as a response to resource scarcity, or as something else entirely? Testing a lens against the data, seeing what it illuminates and what it flattens, is a normal part of ethnographic craft, and it is also the part that manual recoding makes prohibitively slow, since re-tagging a full field-note corpus against a new theoretical frame by hand can take as long as the original coding pass.
Skimle's category restructuring lets you reorganise an already-coded corpus under a different conceptual frame without recoding from the raw text each time, which is the practical mechanism behind criterion 6, method-agnosticism, in our design criteria: a tool should support theoretical reframing rather than lock a study into whatever categories emerged first. Agentic analysis goes a step further for research-question-driven work, applying criterion 7, adversarial rigour, by running a dedicated counter-evidence search against each theme it proposes, so a pattern that looks clean in aggregate gets checked against the cases that do not fit it.
That counter-evidence step matters more in ethnography than the design criteria post's general framing suggests, because language models are drawn toward the frequent and the consensual by default. Researchers led by Valentin Hofmann showed in Nature that models can hold stereotypes about specific groups more extreme than any human stereotype recorded experimentally, which is a sharp warning against letting any AI-generated pattern stand unexamined when a study concerns a specific community or subculture. An ethnographic reframing exercise that only surfaces the dominant pattern and never the deviant case has reproduced exactly the majority-flattening effect our post on bias in AI-assisted qualitative analysis describes. Explicitly asking for what contradicts the emerging pattern, not just what supports it, is the corrective.
What can't AI do in ethnographic research?
The limits matter as much as the capabilities, and they are sharper in ethnography than in most qualitative methods.
Participant observation itself. No tool can build the rapport that lets a researcher into a setting, read a room in real time, or decide which of a hundred simultaneous events is worth writing down. That judgement is exercised in the field, in the moment, and it is the part of ethnography no amount of downstream AI assistance touches.
The 419 signatories to a 2025 article in Qualitative Inquiry, rejecting generative AI for reflexive qualitative research outright, included Virginia Braun and Victoria Clarke, whose thematic analysis framework much of the field rests on. Their argument is explicitly about all phases of reflexive interpretive work, not just initial coding, and ethnography's dependence on the researcher's own positioned, accumulated understanding of a setting makes it a strong candidate for the kind of work they mean. If your ethnographic tradition treats interpretation as constitutively human, and many do, that position should not be argued away by a vendor's feature list, including ours. Our design criteria post names this directly as a limit that better engineering does not resolve.
Thick description cannot be manufactured from a transcript alone. Geertz's point was that the meaning layered onto an observed act depends on context the researcher carries in their head, most of which never gets written down verbatim. AI can help organise and surface patterns in what was written down. It cannot supply the unwritten context that made the written note meaningful in the first place.
Cognitive offloading is a real risk, not a hypothetical one. The literature on heavy AI reliance links it to weaker independent judgement, and the risk is worse in ethnography precisely because so much of the interpretive weight sits outside the text. See is AI destroying the ability to think for the underlying evidence. The discipline that protects against it is the same one that protects any qualitative study: treat AI-surfaced codes and patterns as a draft to argue with, not a finding to accept.
A worked example: a retail ethnography from field notes to write-up
A short example makes the four stages concrete. This kind of workplace or customer-behaviour study is also common outside academia, in UX and market research contexts, so the workflow below applies whether the write-up is heading to a journal or a client deck; commercial teams without an ethics board to answer to should still read responsible AI in qualitative market research alongside it. Suppose you are studying how staff on a supermarket shop floor actually handle customer complaints, as distinct from the official policy in the training manual.
Before fieldwork, you upload a dozen prior ethnographies of retail and service work into a project and run an inductive pass to see which theoretical framings (emotional labour, informal resistance, improvised professionalism) other researchers applied to similar settings, so you go into the field with a sense of the existing conceptual map rather than starting cold.
During four weeks of observation, you dictate field notes into your phone after each shift and record the handful of semi-structured interviews you conduct with staff; both get transcribed and land in the same project as searchable documents, tagged with metadata for store, shift, and week of fieldwork. As the corpus grows, an inductive coding pass surfaces early descriptive categories like "de-escalation scripts" and "off-script workarounds," and you refine, merge, and rename them as you re-read the material against what you actually observed.
By week three, cross-tabulating categories against the week-of-fieldwork metadata shows that "off-script workarounds" becomes far more common after staff have been on shift with the same manager for more than two weeks, a pattern you would not have noticed from memory alone. You then restructure the coded corpus under an emotional-labour framing to see whether that lens explains the pattern better than an informal-resistance framing, running a counter-evidence check against both before deciding which one the data actually supports. The write-up cites the specific field notes and interview passages behind each claim, because every theme in the categories view traces back to its source.
Nothing in that process replaces the weeks spent on the shop floor. It compresses the mechanical layers around that fieldwork, transcription, first-pass coding, and reframing, so more of your time goes toward the interpretive judgement only you can supply. Our comparison of qualitative analysis tools covers how Skimle and the established CAQDAS options (NVivo, MAXQDA, ATLAS.ti) compare if you are weighing which one fits a fieldwork-heavy study.
Frequently asked questions
Can AI replace participant observation in ethnography?
No. Participant observation depends on a researcher's physical presence, real-time judgement about what to attend to, and the rapport that lets them into a setting at all. AI tools can help with the transcription, coding, and pattern-discovery work that follows fieldwork, but the fieldwork itself is not a task AI performs.
How is using AI for ethnography different from using it for interview-based qualitative research?
Interview transcripts are largely self-contained: the meaning is mostly in the words on the page. Field notes depend heavily on context the researcher accumulated over weeks in a setting, which never fully makes it onto the page, so AI-surfaced codes from field notes need more researcher judgement to interpret correctly than AI-surfaced codes from a single interview transcript.
How do I anonymise ethnographic field notes without stripping out the detail that makes the analysis meaningful?
Combinations of descriptors (a specific role, location, and time period together) can re-identify a person or setting even after names are removed, which is a sharper risk in ethnography than in most qualitative work because settings are often unique. Anonymisation that pseudonymises consistently across a corpus, rather than a manual find-and-replace, keeps the analytical detail while reducing that risk. See our IRB-compliant pseudonymisation guide for the practical steps.
Can AI tools analyse photos, video or only written field notes?
Purpose-built qualitative analysis tools, Skimle included, are built primarily around text: transcripts, written field notes, and documents. Photographs and video remain part of an ethnographic record that a researcher reviews directly; where a study relies heavily on visual data, treat AI assistance as covering the written and transcribed portion of the corpus, not the visual material.
How much of an ethnographic study can realistically be AI-assisted without compromising rigour?
The literature review, transcription, and first-pass coding stages can be substantially AI-assisted with proper verification at each step. The interpretive work of deciding what a pattern means, choosing between competing theoretical framings, and the fieldwork itself remain the researcher's. Our 20-question AI qualitative data analysis checklist is a reasonable pre-publication review to run against whatever AI assistance you used.
Want to try this on your own field data? Try Skimle for free and run a batch of field recordings or notes through transcription and an inductive coding pass to see how the workflow fits your study.
Related reading: See 9 design criteria for AI-assisted qualitative analysis tools and designing AI that augments qualitative researchers for the underlying design argument, and reflexivity in qualitative research for how researcher positionality applies to AI-assisted work.
About the authors
Henri Schildt is a Professor of Strategy at Aalto University School of Business and co-founder of Skimle. He has published over a dozen peer-reviewed articles using qualitative methods, including work in Academy of Management Journal, Organisation Science, and Strategic Management Journal. His research focuses on organisational strategy, innovation, and qualitative methodology. Google Scholar profile
Olli Salo is a former Partner at McKinsey & Company where he spent 18 years helping clients understand the markets and themselves, develop winning strategies and improve their operating models. He has done over 1000 client interviews and published over 10 articles on McKinsey.com and beyond. LinkedIn profile
Disclosure: the authors have a financial interest in Skimle.
Sources
- Participant Observation as a Data Collection Method - Kawulich, Forum: Qualitative Social Research (2005)
- The Interpretation of Cultures: Selected Essays - Geertz, C. (1973), Basic Books
- Meta-ethnography 25 years on: challenges and insights for synthesising a large number of qualitative studies - Toye et al. (2014), BMC Medical Research Methodology
- We Reject the Use of Generative Artificial Intelligence for Reflexive Qualitative Research - Jowsey, Braun, Clarke, Lupton and Fine (2025), Qualitative Inquiry
- AI generates covertly racist decisions about people based on their dialect - Hofmann, Kalluri, Jurafsky and King (2024), Nature 633(8028), 147-154



