To the survey: https://docs.google.com/forms/d/e/1FAIpQLSfujaTdkiyrwagGO4Cv_JjGQ_IIWt9v4tO54aVBSiF7J6a_pw/viewform?usp=header
Seeing the first responses come in was a really exciting moment. After spending so much time planning the survey, choosing the images and creating the questionnaire, I was finally able to look at the data. I was curious to find out whether people could actually tell the difference between authentic, AI-edited and fully AI-generated images. Even more importantly, I wanted to understand where people struggled and whether any interesting patterns would appear.
In total, 14 people completed my pilot study. They came from different age groups, although most participants were between 25 and 34 years old. Each participant looked at 24 images and decided whether each one was authentic, AI-edited or fully AI-generated. Altogether, this gave me 336 individual answers to analyse.
The first thing I looked at was the overall recognition rate. After comparing every answer with the correct solution, I found that participants classified only 42.3% of the images correctly. I honestly expected the result to be higher. At first, this seemed surprisingly low, but the more I thought about it, the more it made sense. AI-generated images have become incredibly realistic, and even manipulated images are often very difficult to recognise.
I also wanted to see whether some image categories were easier than others. Authentic photographs achieved the highest recognition rate, with 57.1% of participants identifying them correctly. AI-edited images were recognised correctly in 42.9% of the cases, while fully AI-generated images were the most difficult, with only 26.8% correct answers. This immediately caught my attention because it showed that participants struggled most with images that were created entirely by AI.
While these numbers already gave me a good overview, I quickly realised that there is more to it. Some images performed much better than others, so I decided to analyse every image individually. Looking at the results in more detail helped me understand which images confused participants and which visual characteristics might have influenced their decisions. This became one of my favourite parts of the analysis.
Besides the recognition rates, I also collected confidence ratings. After every image, participants indicated how confident they felt about their answer. I included this question because I wanted to know whether people were aware when they were uncertain. It is one thing to answer correctly, but it is another thing to know how reliable your own judgement actually is. I will analyse this in more detail in the next blog post.
One thing that became clear after analysing the pilot study is that recognising AI-generated images is more difficult than many people might expect. Before starting the study, I already assumed that people tend to overestimate their ability to recognise AI-generated images. Although this pilot study is too small to confirm that assumption, the results encourage me to explore this question further in the main study by comparing participants’ self-assessment with their actual performance.
!!Attention!!
!!Please only look at the documents if you have already completed the study or if you do not intend to participate in it!!