Study title: How R users turn ggplot2 code into reusable functions
Principal investigator: Cynthia A. Huang, Social Data Science and AI Lab, Institut für Statistik, LMU München, Ludwigstr. 33, 4. OG, 80539 München
Contact: cynthia.huang@lmu.de
Data Protection Officer: Dr. jur. Marco Wehling, LL.M., Behördlicher Datenschutzbeauftragter der Ludwig-Maximilians-Universität München, Geschwister-Scholl-Platz 1, D-80539 München. Contact via the official contact form
Purpose and procedure
This study investigates the situations in which people who use ggplot2 consider turning plotting code into reusable helper functions, what they end up writing, and how they reflect on it afterwards. We are interested in your experience regardless of your level of expertise — whether you've written a single one-off function for yourself, decided not to, or maintain a published package, your perspective is useful to this research.
You will be asked to complete an online survey (~7 minutes, up to ~15 minutes if you choose to describe several occasions) covering your background with ggplot2, your general practice with reusable plotting code, and up to three specific occasions on which you wrote or considered writing such code. For each occasion you will be asked to paste code snippets or link to public code (e.g. GitHub) showing what the code looked like before (the original) and after (the solution).
Data collected
| Data type | Purpose |
|---|
| Event/context, prior exposure to the talk or blog post | Sample description; controlling for prior exposure |
| Demographic and role information (e.g. academia/industry, ggplot2 experience, role) | Sample characterisation |
| Multiple-choice and rating-scale responses about specific occasions of writing plotting functions | Core research data |
| Free-text descriptions of those occasions and your reflections on them | Core research data |
| Code snippets or links to public code repositories | Core research data — describing the structure of the original code and the solution |
No sensitive personal data (health, financial, political opinions, etc.) is collected, and no compensation or prize incentive is offered for completing this survey.
Note on code and links. Please only paste code you are entitled to share — do not paste proprietary or confidential code, and do not include credentials, file paths, data values, or comments that identify people. If you cannot share code as-is, you may paraphrase it or use an AI assistant to remove sensitive details before pasting. A link to a public repository may identify you as its author; if you prefer to remain unidentifiable, paste an excerpt with identifying details removed instead of linking. Code excerpts may be quoted in publications in anonymised form (e.g. with variable and function names altered where they would identify a project or organisation). You may skip a code field, but occasions without code may be excluded from analysis.
Note on anonymity in practice. While no single field in this survey directly identifies you, an unusual combination of role, experience level, free-text answers, and code could in principle be recognisable to someone who knows you, especially if you're one of few respondents from a particular event or background. Please keep this in mind when writing free-text answers.
Legal basis for data processing
Your data are processed on the basis of your freely given, informed consent, in accordance with Art. 6(1)(a) GDPR.
Data storage, processing, and retention
Survey responses are collected via SoSci Survey, a third-party survey provider hosted on servers in Germany. SoSci Survey acts as a data processor on our behalf under a GDPR Art. 28 data processing agreement; no data collected here is transferred to any other third party. After initial collection, data is processed only by the research team (LMU-affiliated).
Survey responses are collected anonymously: no participant ID, IP address, or other identifier links a submission to you, so once you submit the survey we cannot locate or delete your specific responses — there is nothing in the data that identifies which submission was yours. If you paste a link to a public repository, that link is treated as part of your anonymous response; we will not contact repository owners.
Anonymised research data will be retained for at least 10 years after publication in line with DFG good scientific practice guidelines, and may be deposited in a public repository in fully de-identified, aggregated form. Code excerpts will only be deposited after identifying names, paths and comments have been removed.
Your rights under GDPR
No personal data that identifies you is collected in this study. Because your survey responses are collected anonymously, we cannot identify which submission is yours, so we are unable to access, correct, restrict, or delete an individual response after it has been submitted (Arts. 15–21 GDPR do not apply to data that cannot be attributed to you, Art. 11 GDPR).
You have the right to lodge a complaint with the supervisory authority: Bayerischer Landesbeauftragter für den Datenschutz (BayLfD) — the competent authority for public bodies such as LMU (not the BayLDA, which covers the private sector). For any questions, contact the PI or the Data Protection Officer above.
Voluntary participation and withdrawal
Participation is entirely voluntary. You may stop at any point before submitting by closing the browser window; incomplete responses are deleted. Because responses are collected anonymously, once you submit the survey we cannot identify or withdraw your specific response — if you want the ability to withdraw after the fact, please do not submit, or withdraw before that point.
Declaration of consent
By selecting "Yes" below, you confirm that you have read and understood the information above, have had the opportunity to ask the research team questions (contact details above), and consent to participate in this study and to the processing of your data as described.