Evidence read
How it was rated
The most interesting number in the Spanish pilot's evaluation is not a high score. It is a gap: the strongest positive result in the paper came from the teachers, and the pupils, rating the same term's work, came out neutral.

Two verdicts from the same classrooms
The evaluation of the Spanish pilot reports one result that is easy to quote and one that is easy to skip past, and both came out of the same two schools in the same two years. Teachers rated the Project Management Dashboard, the web tool they used to set projects and review what pupils handed in, as a desired product that satisfied their expectations: the strongest positive result in the paper. The 115 pupils ratingCreate@School, the app they built the games in, came out neutral on average, and agreed with each other closely, which makes that hard to dismiss as noise.
Before this turns into pupils against teachers, be precise about what was rated. The dashboard was evaluated by teachers only, because it was built for teachers, and an adult's verdict on their own planning tool is not like-for-like with a teenager's verdict on an authoring app. The comparison that is like-for-like sits underneath and points the same way: both groups rated Create@School, teachers accepted it positively, pupils landed on neutral.
- Pupils rating the app 115, of 308 who used the tools
- Educators 16
- Ages 8 to 17
- Schools Two, in Úbeda and Puerto de Santa María
- Duration Two academic years, one per cycle
What the survey actually asked
The instrument was the Hassenzahl model, measured with AttrakDiff surveys. It splits a response to software into four dimensions: pragmatic quality, meaning practicality and functionality; hedonic identity, how far the thing lets you show who you are; hedonic stimulation, whether it offers something new to get better at; and attractiveness, the overall appeal. Each is measured with seven pairs of opposed words on a scale from minus three to plus three, zero being neutral: confusing against clear, typical against original, ugly against beautiful. The diagrams also carry a confidence rectangle, a spread measure where smaller means more agreement. The pupils' was very small.
Now the part that has to be unmistakable. AttrakDiff asks whether people found a tool practical, stimulating, identity-fitting and attractive. It does not ask whether anybody learned more mathematics, or science, or language, and it cannot answer that question. Nothing on this page is evidence about learning outcomes. If what you need to know is whether pupils who built a game understood the topic better than pupils who did not, this is not the study that tells you.
What a neutral rating from a teenager means
Both of the easy conclusions are wrong. It is not a failure. The authors report that Create@School satisfied the pupils' expectations and that the application was accepted, and the word-pair results are positive rather than flat: the app is characterised as practical, creative, connective, presentable, appealing and easy to use. Stimulation and identification sat in the average region between zero and two, attractiveness slightly higher. Neutral here is an average across four dimensions, not a shrug on all of them.
Nor is it a triumph waiting to be spun. Neither Create@School nor Pocket Code reached the maximum rating on any quality, and at that stage neither was yet the desired product for the people using it. The paper says so plainly about its own tools, and a page that drops that sentence in favour of the dashboard result is misreading the evaluation on purpose.
The reasons the authors give
These are more interesting than the scores, and they are the authors' explanations rather than ours. The first is ownership. The apps had been chosen by the school, so pupils treated them as a class resource rather than something they had picked, and the paper names this as one reason for the neutral rating. Anyone who has handed a school-issued tool to a class of fourteen-year-olds will recognise the reaction.
The second is that the network took the app down with it. Connectivity in the pilot schools was poor, and the disconnections that followed were perceived by pupils as a fault of Create@School itself rather than as an external factor, which is what led them to characterise the app as unpredictable and unruly. Because school security did not permit open browsing in class, pupils could not fetch images for their games either, and appraised the app more conservatively as a result. We have pulled that strand apart onthe school network reality.
The third is that the novelty wore off.
The novelty paradox
Pocket Code came first. Pupils and teachers found it novel at the start, and when teachers compared the two apps, Pocket Code scored better on professional look, innovativeness and novelty, which the paper puts down to the impact of first use. By the second cycle, Create@School had become part of everyday classroom activity and the sense of novelty diminished. The authors name the paradox themselves: being accepted as ordinary classroom practice is a sign of success, and it lowers the novelty score at the same time.
That is worth sitting with if you ever evaluate a classroom tool yourself. Asking pupils whether they find something exciting is a fair question in week one and a misleading one in month eight, because by then the honest answer from a well-adopted tool is no. The trajectory you want, a tool that stops being an event and becomes a resource, registers on that instrument as a decline, and an evaluation tracking only novelty will recommend replacing whatever is working.
What we would take into a staffroom
Everything above comes from the evaluation. What follows is ours: classroom experience, not the paper, so weigh it differently.
Ask twice, early and late, and expect the second answer to be lower. One survey at the end of a term cannot tell you whether a low stimulation score means the tool is dull or means it has become furniture, and those need different responses from you.
Separate the tool from the infrastructure in the question you ask. The pupils here could not make that distinction, and the evaluation does not let us work out how much of the neutral rating was really the Wi-Fi. Asking what exactly stopped working, at the moment it stopped, gets you closer than a general rating and protects a decent tool from taking the blame for a bad network.
Treat neutral from a class that did not choose the tool as a reasonable ceiling rather than a disappointment. Our bar for anything a school issues is that pupils use it without resistance and make something they will show other people, and by the authors' account of what pupils valued, building a game their friends could play, that bar was cleared. The dashboard says the same from the other side: what made it land was quantified feedback teachers did not otherwise have, and its one named weakness was that grades could not pass to the school management system.
How far this goes
This is 115 pupils in two Andalusian schools between 2015 and 2017, on the tablets and school networks of that time, and it is worth reading alongsidewhat the trial set out to test. The paper also reports the Spanish pilot only: three separate evaluation studies were planned, one per pilot site, and we have not located the British or Austrian ones. So none of this is the project's verdict. It is one country's, honestly reported, which is more than most classroom tools get.
Source. Gaeta, E.; Beltrán-Jaunsaras, M. E.; Cea, G.; Spieler, B.; Burton, A.; García-Betances, R. I.; Cabrera-Umpiérrez, M. F.; Brown, D.; Boulton, H.; Arredondo Waldmeyer, M. T. “Evaluation of the Create@School Game-Based Learning–Teaching Approach.” Sensors 2019, 19(15), 3251.doi:10.3390/s19153251
Open access under CC BY 4.0. MDPI blocks a good deal of automated traffic, so if the DOI link will not open for you, the full text is also at Europe PMC:PMC6695907.