top of page

Projects 

Impresso II: Media monitoring of the past II

Impresso is an interdisciplinary research project that uses machine learning to pursue a paradigm shift in the processing, semantic enrichment, representation, exploration, and study of historical media across modalities, time, languages, and national borders.

Impresso Datalab - Programmatic access to Impresso's Corpus, Data and Models

As managing editor of the Impresso Datalab, I'm responsible for designing and implementing an editorial process to create, review and publish Jupyter notebooks to foster use of Impresso data and models. I also contribute to the design of the interface, user experience, development and documentation of the Python library and models. Besides, I'm responsible for outreaching activities, developing and delivering training for potential new users.  

Case-study: Historical media coverage of the Olympics Games

 

By employing computational methods to analyse news content, I seek to understand how media narratives of the Olympic Games are produced and reproduced. I've previously worked on the concept of Olympic legacy as narrated by the British and Brazilian news media regarding the 2012 London and 2016 Rio Games. Now, I am investigating the relationship between mega-events and cities as narrated in digitised historical newspapers. 

Project's documentation

Impresso Datalab | Webpage of the Impresso Datalab 

Datalab editorial pipeline | Documentation of the editorial pipeline developed and implemented for the Impresso Datalab

Paper | Shaping History, Responsibly: Seven Principles to Guide the Design of Archival AI Assistants for Cultural Heritage Collections

Screenshot 2025-08-20 at 15.21.35.png

This project has received funding from The Swiss National Science Foundation (SNSF) and The Luxembourg National Research Fund (FNR). 

Workshops 

CAO, V., Guido, D., Mello, C., Affan, F., & During, M. (2026, June 16). Impresso: (Re)Thinking Historical Sources and Collections Retrieved with AI Tools in Large Digital Infrastructures [Workshop]. AI through History, History through AI conference 2026, Luxembourg. https://www.uni.lu/c2dh-en/events/ai-through-history-history-through-ai/#tuesday-16-june-2026

CAO, V., Mello, C., Guido, D., Affan, F., & During, M. (2026, July 29). Impresso Datalab: Embedding Newspapers and Radio Archives for Multimodal and Multilingual Data Analysis [Workshop]. DH2026, Daejeon. https://dh2026.adho.org/

Düring, M., Beelen, K., & Mello, C. (2025, June 3). Impresso Datalab Workshop: Programmatic Access and Annotation Services for Multilingual and Multimodal Historical Media Collections [Workshop]. DH Benelux 2025, Amsterdam, Netherlands. https://zenodo.org/records/15437478

Düring, M., Cao, V., Mello, C., Guido, D., & Affan, F. (2025, November). AI and media collections: Impresso Web App and Datalab [Workshop]. Assises thématiques sur l’intelligence artificielle, Differdange, Luxembourg.

Düring, M., Mello, C., Beelen, K., Guido, D., & Ehrmann, M. (2025, July 14). Impresso Datalab Hackathon: Programmatic Access and Annotation Services for Multilingual and Multimodal Historical Media Collections [Workshop]. DH2025, Lisbon.

Düring, M., Mello, C., & Cao, V. (2025). Impresso Datalab: Exploring and annotating multilingual digitised newspapers and radio archives [Workshop]. Sixth Conference on Computational Humanities Research, Luxembourg.

Düring, M., Mello, C., Guido, D., & Kalyakin, R. (2025, March 3). Towards interoperability: Introducing the Impresso data lab for the enrichment and analysis of historical media [Workshop]. Digital Humanities im deutschsprachigen Raum, Bielefeld, Germany. https://zenodo.org/records/15269143

During, M., Vy, C., & Mello, C. (2026, June 2). Impresso Datalab: Embedding Newspapers for Multimodal and Multilingual Data Analysis [Workshop]. DH Benelux 2026, Maastricht. https://doi.org/10.5281/zenodo.19661927

Conference presentations

During, M., CAO, V., Beelen, K., Mello, C., Ehrmann, M., Clematide, S., Boros, E., Conti, P., Michail, A., Opitz, J., Grandjean, M., Michelet, A., Affan, F., Ruppen Coutaz, R., Guido, D., & Mitsurov, K. (2026, July 29). Meaning, Similarity, and Relevance. Reflections on the Integration of Multilingual and Multimodal Embeddings in Historical Media Research [Long paper]. DH2026, Daejeon. https://doi.org/10.5281/zenodo.21495909

Finn, F., Mello, C., & Maurer, Y. (2025, December 5). Developing Archival AI chatbots: Risks, benefits, and future directions. In Human-Centred AI / the UX of AI, AI4LAM Fantastic Futures [Long paper]. https://doi.org/10.23636/54k0-ny43

Mello, C. (2026a, June 26). Introducing Impresso at Co-Creating Digital Futures: Building Bridges Between Cultural Heritage, Open Knowledge, and Generative AI [Round table]. The Digital Conference 2026 @ King’s College London, London.

Mello, C. (2026b, June 27). Exploring Digitised Newspapers Collections with Computational tools. In Exploring colonial histories through digital archives: An introduction to Impresso and computational methods for Historians [Long paper]. 50th Annual Meeting of the French Colonial Historical Society (FCHS), Maynooth. https://frenchcolonial.org/2026-50th-meeting-maynooth/

Mello, C., Finn, F., Khosrowi, D., During, M., & Guido, D. (2026, June 23). Designing archival AI chatbots for enhanced research methods: An introduction to the Impresso Barista. In Decision Systems and Anthropomorphisation of AI [Long paper]. The Digital Conference 2026 @ King’s College London, London. https://www.kcldigitalconference.com/

Jupyter notebooks

Boros, E., & Mello, C. (2025). Impresso Datalab—News Agencies Recognition and Linking with Impresso BERT models [Computer software]. Zenodo. https://doi.org/10.5281/ZENODO.17227166

Kalyakin, R., Düring, M., & Mello, C. (2025). Impresso Datalab—Visualising Place Entities on Maps [Computer software]. Zenodo. https://doi.org/10.5281/ZENODO.17227295

Kalyakin, R., Ehrmann, M., Mello, C., & Düring, M. (2025). Impresso Datalab—Introduction to the Impresso Python Library [Computer software]. Zenodo. https://doi.org/10.5281/ZENODO.17184962

Kalyakin, R., & Mello, C. (2025a). Impresso Datalab—A quick guide to searching with Impresso library [Computer software]. Zenodo. https://doi.org/10.5281/ZENODO.17227217

Kalyakin, R., & Mello, C. (2025b). Impresso Datalab—Exploring Entity Co-occurrence Networks [Computer software]. Zenodo. https://doi.org/10.5281/ZENODO.17227225

Mello, C., Conti, P., Grandjean, M., & CAO, V. (2025). Impresso Datalab—Inspecting my collection with data visualisation tools [Computer software]. Zenodo. https://doi.org/10.5281/ZENODO.17227057

Vinarskis, G., & Mello, C. (2025a). Impresso Datalab—Language Identification with impresso-pipelines Package [Computer software]. Zenodo. https://doi.org/10.5281/ZENODO.17227055

Vinarskis, G., & Mello, C. (2025b). Impresso Datalab—OCR Quality Assessment with impresso-pipelines Package [Computer software]. Zenodo. https://doi.org/10.5281/ZENODO.17225993

Programming Historian

The Programming Historian journal publishes peer-reviewed lessons on digital methods for non-specialist readers. I collaborate with the English journal as editor. I have also previously contributed as reviewer and translator for the PH in Portuguese. 

Project's documentation

Edited lessons

Reviewed lessons

Translated lessons 

woman-using-tabulator.png

Sentiment in the news

Center for Advanced Internet Studies Fellowship programme (2023)

Interpreting sentiment analysis outputs to categorise emotions in news articles

 

Sentiment analysis (SA) is one of the techniques commonly used in Natural Language Processing. These algorithms classify words on a scale of sentiment that usually goes from negative to positive. These techniques have been mostly applied to Twitter data or customer feedback to track reception of products. However, researchers in the humanities have been using the algorithms to explore other kinds of text, such as news articles to understand the tone in which some media events are narrated.

 

This kind of use of the SA presents several challenges. Questions arise about the extent to which the SA scores are reliable or how they can be used for interpretation (qualitative analysis) and discourse analysis. This project aims to investigate how these techniques have been applied to the study of media events, identifying their limitations and potentialities, making use of explainable AI and qualitative analysis of SA outputs in a multilingual text corpus.

 

newssent.jpg

This project has received funding from CAIS via the Ministerium für Kultur und Wissenschaft des Landes Nordrhein-Westfalen.

Past projects 

Discourse analysis of Social Media regulation debates

This project investigates the complex debates on regulatory frameworks for social media platforms in Germany, the United States, and Brazil. We look, in particular, at discourses on the challenges of regulating online speech in a way that respects democratic values, taking into account cultural-specific sensitivities.

 

Germany’s NetzDG reflects an European regulatory model which has raised concerns about ‘overblocking’, contrasting with the United States’ regulatory model based on Section 230 CDA, which champions platform autonomy and freedom of speech. In Brazil, the debate around the Marco Civil da Internet and current social media disputes underscore unique regulatory challenges in Latin America, where recent violent events have intensified calls for platform accountability.

 

By conducting discourse analysis on various media sources, this project compares public narratives on platform regulation, identifying key stakeholders, arguments, and communicative strategies. The analysis seeks to reveal how historical and cultural contexts shape national approaches to online speech regulation.

Project's documentation

Conference Paper | Linked Resources in Debates About the German Network Enforcement Act on Twitter

Funding | Led by Dr. Jens Pohlmann and I, this project has been awarded with a Working Group grant from the Center for Advanced Internet Studies (CAIS)

 

newssent.jpg

This project has received funding from CAIS via the Ministerium für Kultur und Wissenschaft des Landes Nordrhein-Westfalen.

CLEOPATRA

The CLEOPATRA ITN, a Marie Skłodowska-Curie Innovative Training Network led by the L3S Research Center at the Gottfried Wilhelm Leibniz University of Hannover, aims to make sense of the massive digital coverage generated by the intense disruption in Europe over the past decade – including appalling terrorist incidents and the dramatic movement of refugees and economic migrants.

Nationalism, internationalism and sporting identity: the London and Rio Olympics

My research explored the media coverage of the Olympic Games in a cross-cultural, cross-lingual and temporal perspective. I was especially interested in comparing how the concept of ‘Olympic legacy’ has been approached by the Brazilian and British media considering different locations, languages and social-political contexts.

Project's documentation

1. The archived Olympics: exploring the narratives of Legacy in the UK Web Archive

 

Blog posts

 

Podcast - Sport in History Podcast by the British Society of Sports History

Episode: Documenting the Olympics - GLAM sector and London 2012

Recording of the event organised by the British Library in collaboration with the British Society of Sports History (BSSH), the International Centre for Sports History and Culture at De Montfort University (ICSHC) and the School of Advanced Study (SAS)/CLEOPATRA project.


 

 

 

Videos 

 

UK Web Archive conference 2022 - An event about web archiving aimed at staff from UK Legal Deposit Libraries.

 

​Engaging with Web Archives Conference 2020

2. Unveiling the potentials and limitations of Sentiment analysis algorithms to study news media

Paper | Combining sentiment analysis classifiers to explore multilingual news articles covering the London 2012 and Rio 2016 Olympics

Clo-poster2.jpg

This project has received funding from the European Union’s Horizon 2020 research and innovation programme under the Marie Skłodowska-Curie grant agreement no. 812997.

#OccupyEstelita: the emergence of identity as part of a political performativity in the usage of Facebook Events by Social Movements

Many people who joined the Spanish 15-M, Occupy Wall Street and the Brazilian #OccupyEstelita used to recuse the word activist, arguing that this word could not represent their activities: virtual participation supporting Facebook Events, sharing information about the movement online and inviting virtual friends to be a part of the demonstrations. This project investigates the following question: How is the perception of a political self-identity affected by the usage of Social Networks? The main subject of study will be the usage of Facebook Events by manifestants from the Occupy Estelita movement in Brazil. The project has conducted in-depth interviews and questionnaires with movement’s supporters to understand how they perceive their online activities as part of a social mobilization aimed at physical presence on the street.

Project's documentation

CAIS Fellowship Report | #OccupyEstelita: The Emergence of Identity as Part of Political Performativity in the Use of Facebook Events by Social Movements  

Paper | Facebook Event as a platform to promote engagement in social movements: Theory of performativity applied to social networks

Book Chapter (in Portuguese) | Performativity and conflicts between the occupation of real and virtual spaces: a study of the Facebook Events platform

imagem_2.webp

This project has received funding from CAIS via the Ministerium für Kultur und Wissenschaft des Landes Nordrhein-Westfalen.

bottom of page