Practice report
The Spanish pilot schools
Two ordinary state schools in Andalusia spent two academic years building games inside normal subject lessons. There is a published evaluation of what happened. It covers Spain only, and it measures whether people accepted the tools, not whether anyone learned more.

What the trial asked of two Andalusian schools
The Spanish end of No One Left Behind ran in two state schools, one in Úbeda and one in Puerto de Santa María, across two consecutive academic years between 2015 and 2017. It was not a coding club and not an IT option. Pupils built games inside subjects the timetable already held: science, mathematics, biology, geology, language, social sciences and computing, plus enrichment, a multidisciplinary programme available in Andalusia, and PEMAR, a programme for children with attention deficit disorders.
The two years ran as two cycles with different tools. In the first, classes usedPocket Code, the mobile block-based environment from TU Graz's Catrobat project, in a version that predated the school edition; its job that year was to train pupils and teachers and to produce a first set of templates. In the second,Create@School and the teachers' Project Management Dashboard were completed and validated in lessons. So the pilot did not test a finished product. It built one while teaching with it, which is worth holding on to when you read the ratings.
What this evaluation covers, and what it does not
Everything below comes from one paper, and that paper reports Spain. Because the three pilot sites had different contexts, three separate evaluation studies were planned, one each for Spain, the United Kingdom and Austria, and this is the Spanish one. We have not located the British or the Austrian study. Readers who care about the Graz end of the project should know that nothing here speaks to it.
The second limit matters more. The instrument was the Hassenzahl model with AttrakDiff surveys, which ask whether people found something practical, stimulating, identity-fitting and attractive. That is acceptance and user experience, not attainment. Nothing here shows that pupils learned more mathematics than they would have otherwise. The authors do conclude that programming principles promoted logical thinking and stimulated creativity, but that is their characterisation of what they saw, not a learning result the study measured. If someone has sent you this paper as proof that game-making raises results, it is not that paper.
Who took part
- Where Two schools in Andalusia: Úbeda and Puerto de Santa María
- When Two academic years, one per cycle, between 2015 and 2017
- Pupils 308 used and validated the tools; 115 rated the experience
- Ages 8 to 17, 6th to 11th grade
- Educators 16
- Devices 338 tablets and mobile phones
- Tools Pocket Code in cycle 1; Create@School and the teacher dashboard in cycle 2
- Measured Acceptance and user experience, not learning outcomes
- Planned project scale ~600 pupils, 3 countries; 5 schools per CORDIS, 8 per the evaluation
The two pupil numbers are not interchangeable. 308 pupils used and technically validated the tools across both cycles; 115 filled in the surveys that produced the ratings. When you read below that pupils rated the app neutral, that is those 115. The split is printed as about 45% girls and 54% boys, which sums to 99, and we are not going to guess at the missing point. Sixteen educators took part: six with only the preliminary Pocket Code version, six with only Create@School, four with both. Parents gave informed consent and were briefed beforehand.
The last row of the box needs a note. CORDIS puts the planned project at some 600 children across 9 to 12 subjects in five schools; the evaluation describes the same two-cycle experiment as reaching over 600 students from eight different schools. Both are project-level statements about the same project, and we are not going to quietly pick one. For the Spanish findings it changes nothing: this pilot was two schools.
The two validation cycles of the Spanish pilot
Cycle one ran in the first academic year with Pocket Code across four subjects: science, enrichment, mathematics and the PEMAR programme. Cycle two ran in the second academic year with Create@School and the teacher dashboard across nine subjects: computing, mathematics, science, programming basics, biology, geology, language, social sciences and enrichment.
1First cycle
First academic year
Pocket Code, as the pre-design version
4 subjects
- Science
- Enrichment
- Mathematics
- PEMAR programme
2Second cycle
Second academic year
Create@School and the teacher dashboard
9 subjects
- Computing
- Mathematics
- Science
- Programming basics
- Biology
- Geology
- Language
- Social sciences
- Enrichment
The tools in the room
338 tablets and mobile devices, seven-inch and 10-inch Android models such as the Google Nexus 7 and the BQ Edison 3. Modest hardware, deliberately: kit a school might plausibly own rather than a computer suite. Their sensors are why this work ended up in a sensors journal: a racing template only works because accelerometer and gyroscope data can be reached from a block a twelve-year-old drags into place. What Create@School added was pre-coded game templates, per-pupil accessibility settings through the GPII framework, and the dashboard, where teachers assign projects, collect uploads and mark them against curricular objectives.
What came back
Three results carry the weight. Teachers rated the dashboard a desired product, the strongest positive finding in the paper. Pupils rated Create@School neutral on average, and agreed closely with each other. And the templates saved more than 40% of coding time, which is the reason the authors give for Create@School beating Pocket Code on pragmatic quality with teachers. The gap between the teacher verdict and the pupil verdict has a page of its own:how it was rated works through the four dimensions.
Neutral is not rejected. The paper's own summary is that the app satisfied pupils' expectations even if it was not yet their desired product, and that it was still accepted. Two of the reasons offered for the flat score have nothing to do with the software. The apps had been chosen by the school, so pupils treated them as a class resource rather than something of their own. And the novelty wore off: what felt new with Pocket Code in year one was ordinary classroom practice by year two. The authors name the paradox themselves, since becoming unremarkable is what adoption looks like and it costs points on a novelty scale.
Then there is the infrastructure. Poor classroom Wi-Fi, authentication rules that blocked browsing during lessons, and a network that could not carry that many devices at once produced disconnections which pupils read as faults in the app rather than faults in the building. That misattribution is why the app was called unpredictable and unruly. Those findings generalise further than anything else in the paper, so they have their own page:what the school network did to the pilot.
One strand ran outside lessons. The Úbeda school used both apps for the LEGO League, building a robot that separated organic from non-organic waste, controlled with advanced blocks and the phone's own sensors alongside the LEGO ones. The team came second in the local championship of Granada and, in the provincial league that followed, took the robot designers and young promise awards. More on that in the LEGO League robot.
What a school should take from this in 2026
Two things here are worth more to a school today than the ratings are. The first is the 40% figure. Starting pupils from an almost finished game rather than an empty screen does not depend on which app you use, and it is the difference between a period that produces something and a period that produces a login screen and a bell. That argument is set out inwhy half-finished games work better. The second is the dashboard's one named weakness: it did not talk to the school management system, so grades could not transfer, and the paper says plainly that this reduced its appeal. That is procurement rather than pedagogy, and it decides whether a tool survives its second term.
Our own judgement, separate from the paper: the pupil rating is the least alarming number here and the network finding is the most useful one. A neutral score from 115 teenagers about software their school chose is about what you would expect. A network that drops a class of 25 mid-task will sink any tool you put in front of them, and then get blamed on the tool. Test the room before choosing the app: 25 devices, one period, everyone uploading in the last five minutes.
One caveat: Create@School was a project deliverable and is no longer operated, while the Catrobat and Pocket Code work continues at TU Graz. The approach transfers, the app does not. The classroom hub has the rest.
Sources. European Commission, CORDIS:project 645215.
Gaeta, E.; Beltrán-Jaunsaras, M. E.; Cea, G.; Spieler, B.; Burton, A.; García-Betances, R. I.; Cabrera-Umpiérrez, M. F.; Brown, D.; Boulton, H.; Arredondo Waldmeyer, M. T. “Evaluation of the Create@School Game-Based Learning–Teaching Approach.” Sensors 2019, 19(15), 3251.doi:10.3390/s19153251. Open access under CC BY 4.0; free full text atEurope PMC PMC6695907.