Аннотация
Language models are increasingly proposed as assistants for qualitative coding, but reliability evidence remains scarce. Two human coders and three model configurations independently coded 600 interview excerpts against an established scheme. Model-human agreement was substantial for descriptive codes and poor for interpretive ones, where models systematically over-applied the most frequent category. Agreement also degraded across the corpus as excerpt length increased. We recommend restricting automated coding to descriptive passes with human adjudication retained throughout.