An open survey dataset on study habits and AI use among university students: A proportionally sampled multi-program study
Jessica María Rojas-Mora, Diógenes de Jesús Ramírez-Ramírez, Cristian David Correa Álvarez
Well-documented survey datasets are still scarce for examining how study habits and generative artificial intelligence (AI) intersect in higher education, especially in Latin America. This Data Descriptor aims to document and enable reuse of a cross-sectional dataset on study habits, academic self-perception, AI use, and ethical perceptions among undergraduate students. Data were collected at the Manizales campus of a Colombian public university from February 15 through May 15, 2025, during the first academic term of 2025. The sampling frame comprised 4,950 undergraduates across 14 academic programs, and the released dataset includes 357 complete responses and 29 survey variables collected through proportional program allocation and quota-based recruitment. The instrument covers sociodemographic and academic characteristics, study routines, AI tool use, perceived usefulness, verification practices, academic integrity, dependence, creativity, and prompt engineering knowledge. The repository provides an anonymized raw dataset, a cleaned analytical dataset, the sampling frame, a codebook, a bilingual questionnaire, and a reproducible R script. Technical validation checks address structural integrity, range plausibility, cross-file reconciliation, and semantic harmonization of AI-tool entries.