Back to Programme

Safeguarding the Scientific Integrity of Online Studies in the Era of AI

Training workshop – Monday, October 26, 13:00-16:00
Language of instruction: English

Online studies are a key data collection tool in the social sciences and beyond. However, because of novel respondent recruitment strategies, including social media platforms and river sampling, these studies are increasingly threatened by bots (i.e., programs that autonomously interact with systems, such as online studies). Bots range from simple rule-based applications to sophisticated applications linked to Large Language Models (LLMs). In particular, this workshop will introduce the challenges and opportunities of LLM-based bots in the context of online studies, focusing on their technical foundations, impact on data quality and integrity, and potential use for online study pretesting. This will be accompanied by an outlook on novel research avenues that go beyond bots, including considerations on scientific integrity and research ethics.

The workshop will begin with an introduction to recent developments in online study recruitment that have made online studies vulnerable to bot infiltration. You will learn what bots are, how they operate, and how they have evolved alongside advances in AI. Building on this foundation, the course will provide hands-on training in programming online studies, including bot prevention and detection methods. Subsequent workshop units will focus on the technical aspects of rule-based and LLM-based bots, including web scraping basics, prompt engineering, persona prompting, and response synthesis.

The workshop will also concern the collection of paradata, such as response times and keystrokes, to detect bot-specific completion behaviors in online studies. In addition, we will demonstrate how to utilize bots for efficient online study pretesting. In doing so, we will switch our focus on their methodological merits to, for example, test complex online study designs. Finally, the workshop will conclude with a discussion of novel research avenues as well as ethical considerations and legal regulations (e.g., EU AI Act).

The workshop is organized into five units, including hands-on exercises. For example, I will provide tailored input on bots and their methodological scope as well as on developing and evaluating bot prompts, implementing bot detection features, cleaning web survey data, and working with bot-generated pretesting data (just to name a few).

The course will be useful to students, researchers, or professionals who are working with online studies as well as those interested in data quality and integrity. By the end of the course, you will understand recent developments in online study recruitment; know how to effectively protect online studies against LLM-based bots; know how to streamline online study pretests through bots; anticipate novel research avenues that go beyond bots; understand bots in light of scientific integrity and research ethics.

Requirements: basic knowledge in online study programming as well as data processing and analysis. Welcome, but not mandatory: basic knowledge in script/programming languages (e.g., HTML and JavaScript). During the workshop, it is recommended that participants have access to a chat-based general-purpose Large Language Model (LLM), such as OpenAI’s ChatGPT or Google’s Gemini. In addition, it would be good to open spread sheets, such as those of Microsoft’s Excel.

Instructor

Jan Karem Höhne is Professor of Survey Methodology at Leibniz University Hannover and Head of the CS3 Lab for Computational Survey and Social Science at the German Centre for Higher Education Research and Science Studies (DZHW). His research combines survey methodology and computational social science, with particular interests in digital data collection, artificial intelligence, speech and language technologies, and innovative approaches to online surveys. He has held research and visiting positions at institutions including the University of Mannheim, GESIS, Utrecht University, the University of Michigan, and Stanford University, and has published extensively in leading international journals in survey research and computational social science.