Annex 1

Privacy Notice for the CLICOPRE Website

The CLICOPRE project is carried out by the University of Padua, with Harvard University as co-host, and is funded under the Seal of Excellence@UNIPD (SoE) scheme. This notice explains how personal data may be processed within the project and is published for transparency purposes on the project website.

1. Data Controller

The University of Padua is the data controller for the personal data processed within the project.

2. Contact Details

Data Protection Officer: privacy@unipd.it.

Principal Investigator: Dr Duccio Gamannossi degl’Innocenti, duccio.gamannossi@unipd.it.

CLICOPRE project team (for data-subject requests): info@clicopre.com.

3. Purpose of the Processing

CLICOPRE studies public debate on climate-related issues in the European Union and the United States. In particular, the project examines how extreme weather events shape public debate, how audiences react, and whether these shocks translate into legislative action and voting behaviour.

4. Categories of Data Processed

The project relies on three broad categories of data: social media data, third-party datasets, and project-generated data and documentation.

4.1 Social Media Data

The project processes publicly available data from public Facebook pages and public Instagram accounts. Data were collected through CrowdTangle, a tool provided by Meta to approved researchers for access to public content on its platforms. The project complies with Meta’s terms, conditions, and privacy policies.

Accessible fields may include the numeric ID and name associated with the publishing account, the date and time of publication, the post type, attached links, the post ID, the textual description, and aggregated engagement counts such as likes, comments, shares, and reactions. No personal data of individual users who react to or comment on posts are collected through CrowdTangle. However, personal data may still be present indirectly where account names, post content, timestamps, or similar fields make an individual identifiable, whether alone or in combination with other information. In some cases, special categories of personal data, including political opinions, may also be incidentally present in public political communication.

4.2 Third-Party Datasets

The project also uses external datasets to link communication patterns to climate events, policymaking, and political behavior. These datasets do not raise concerns relating to personal data or special categories of data. Where third-party datasets are subject to license or access restrictions, they are handled in accordance with the applicable terms and are not redistributed unless explicitly permitted.

4.3 Project-Generated Data and Documentation

The project also produces cleaned and harmonized datasets, merged analytical files, topic-model outputs, natural-language-processing-derived variables, legislative text indicators, metadata, codebooks, and reproducibility scripts. These internal research outputs are documented and managed to support research integrity, internal reuse, and reproducibility.

6. How the Data Are Collected and Processed

Data collection was carried out through official tools such as Meta CrowdTangle. The project processes only data that is necessary for the approved research objectives. The analysis may include frequency analysis, audience engagement analysis, topic analysis, sentiment analysis, stance analysis, and causal inference techniques. The processing does not involve automated decision-making, including profiling, that produces legal effects concerning data subjects or significantly affects them (Article 22 GDPR).

AI-assisted tools, including large language models, may be used for data validation and content analysis within the scope of the CLICOPRE project, including analysis of the text of posts. The default workflow for post-text analysis is to use open-source models run on the University's own infrastructure within the EEA. OpenAI ChatGPT may also be used under the CRUI-OpenAI agreement for validation and content-analysis tasks, including where inputs may contain personal data or special categories of data such as political opinions, only where the institutional DPA is executed and the workspace is covered by European data residency and inference residency for in-scope customer content. OpenAI Ireland Limited acts as a data processor pursuant to Article 28 GDPR and does not use project data to train its models. Any residual processing outside the European Economic Area is governed by the Standard Contractual Clauses (Article 46(2)(c) GDPR) incorporated in OpenAI's Data Processing Addendum. OpenAI API falls under the same regulations. The AI-assisted tools used in the project do not fall within the category of high-risk AI systems under Article 6 and Annex III of Regulation (EU) 2024/1689. Their use is therefore expected to entail only transparency and information obligations for the deployer.

7. Recipients and Disclosure

Personal data may be accessed only by authorized members of the research team and, where strictly necessary, by authorized research collaborators within the EEA or service providers acting within the limits of the project and subject to appropriate safeguards. Non-anonymized data are not communicated or disseminated publicly. Results are presented only in aggregated or anonymized form. Before publication, any granular information that may lead to the identification of an individual is removed. Platform data is therefore of closed access. Results derived from third-party licensed data are handled in accordance with the applicable license terms and shared only in forms that do not reproduce restricted content. Where AI-assisted analysis is performed under the conditions described in this notice, OpenAI (OpenAI Ireland Limited) acts as a data processor and may process personal data for that purpose, with residual processing outside the EEA governed by the Standard Contractual Clauses. GitHub and Harvard University do not receive personal data in the project workflow.

8. Retention and Storage

The following data categories are considered to have long-term value and will therefore be retained:

  • Aggregated and de-identified datasets. These represent the main public-facing research outputs of the project and will be preserved in a public repository with persistent identifiers.
  • Analysis code, scripts, and computational workflows, together with codebooks, metadata, and related documentation. These materials are necessary to support transparency and reproducibility and will be preserved on GitHub and, where appropriate, in a public repository with persistent identifiers. GitHub hosts only analysis code and documentation.
  • Raw platform-derived data and cleaned or harmonized internal datasets in non-anonymized form. These materials are necessary for verification, research integrity, and compliance purposes. A data sharing agreement may be required for any restricted materials shared for these purposes. The specific terms will be determined on a case-by-case basis. These datasets will therefore be retained in secure storage for as long as necessary for the research purposes, including verification, reproducibility, and research integrity, in accordance with Article 89 GDPR and subject to periodic review at least annually and at major project milestones. The personal data is kept in pseudonymized form rather than anonymized at a fixed point, while anonymized or aggregated outputs are derived from it for sharing and publication. At each review, the PI will assess whether continued retention remains necessary and whether further aggregation or anonymization can be performed.

Data are stored through a layered combination of researcher-controlled and institutional solutions at the University of Padua and Harvard University. Storage and backup arrangements are designed for large files and computationally intensive workflows and include encrypted drives, institutional cloud storage, and secure computing environments with access controls. Personal data is stored primarily within the University of Padua's secure infrastructure in the EEA. OpenAI may process personal data only under the safeguards described in Section 6 of this notice and Section 3.3 of the DMP. The Harvard infrastructure is used only for anonymized or aggregated data, and personal data is not transferred to Harvard.

9. Security Measures

Appropriate technical and organizational measures are implemented to safeguard the rights and freedoms of data subjects. Access to non-anonymized data is restricted to the PI and, only where strictly necessary, to authorized collaborators within the EEA acting under specific instructions and confidentiality obligations. Access to storage and computing environments is protected through institutional accounts, encrypted connections, and project-level access controls.

Non-anonymized and special-category social media data are not transmitted through insecure channels. Non-anonymized data is shared only with authorized collaborators within the EEA and with processors operating under the safeguards described in this notice. Any transfer outside the EEA is limited to anonymized or aggregated data, except for residual processing by OpenAI where covered by the Data Processing Addendum and the Standard Contractual Clauses. Security procedures will be reviewed if project workflows materially change.

Identifiable fields are pseudonymized using a keyed hash (HMAC-SHA-512) with a secret key held only by the PI and kept separate from the data, so that they cannot be re-identified by third parties. Data that is shared externally or published is aggregated or otherwise de-identified so that it does not constitute personal data.

10. Information to Data Subjects and Limits to Rights

The research does not collect personal data directly from data subjects. The University of Padua will respond to data-subject rights requests in accordance with the GDPR, subject to the exemptions and limitations applicable to processing for scientific research purposes.

In line with Article 14(5)(b) GDPR, the University of Padua will not provide individual notice to each data subject regarding the processing, since the data are not collected directly from the data subjects and doing so would require a disproportionate effort. This notice is therefore published on the project website to ensure transparency.

11. Data Subject Rights

Under the GDPR, data subjects may have rights including access, rectification, erasure, restriction of processing, objection, and the right to lodge a complaint with a supervisory authority. The exercise of some of these rights may be limited where the GDPR provides a specific exemption applicable to processing for scientific research purposes. To exercise these rights, data subjects may contact the CLICOPRE project team at info@clicopre.com

12. Right to Lodge a Complaint

Data subjects have the right to lodge a complaint with the competent supervisory authority.

13. Website Use and Statistics

The CLICOPRE website (clicopre.com) is hosted on GitHub Pages (GitHub, Inc., a Microsoft company). To deliver the pages, GitHub processes technical data such as the visitor’s IP address and browser information in its server logs. GitHub is based in the United States, under its own data-protection safeguards. This processing relies on the University’s legitimate interest in operating and securing the website (Article 6(1)(f) GDPR).

The website stores the visitor’s language, colour theme, and last scroll position locally in the browser. This information stays on the visitor’s device, is not transmitted to the University, and can be cleared at any time from the browser. It is technically necessary to provide the chosen features and requires no consent.

Visit statistics are limited to aggregate counts provided by GitHub’s repository traffic feature. No analytics script runs in the visitor’s browser, no cookies or persistent identifiers are used, and the site loads no third-party content. The website performs no profiling, tracking, or advertising. For these reasons no cookie banner is required.