- Reading time: 3 minutes
- Price: Free download
- Published: 4th October 2026
- Word count: 628 words
- File format: Text
Personal statement example
Every night at about two o'clock, the ticket machines from our depot's buses finish uploading their day's transactions to a server I help look after. Most nights the files arrive in order. Some nights a bus that broke down on the ring road reports twelve hours late, and the next morning's passenger figures are quietly wrong until someone notices. For the past two years, as a data support analyst at a regional bus operator, that someone has often been me. Working out how to handle data that arrives late, duplicated or out of sequence has become the problem I most want to study properly, and it is why I am applying for postgraduate study in big data systems.
My undergraduate degree in Computer Science gave me a sound base in algorithms, databases and operating systems, and I graduated with a 2:1. The module I enjoyed most covered distributed systems, partly because it was the first time I saw correctness depend on timing rather than just logic. For my final-year project I built an Apache Spark pipeline to analyse three years of publicly released bike-share trip records from a UK city, around forty million rows. My first version ran on a single machine and took most of an afternoon, mainly because one join shuffled the entire dataset. Partitioning the trips by start date and station, and broadcasting the small station table rather than joining it conventionally, brought the run down to under twenty minutes on the same hardware. The analysis itself, comparing weekday and weekend usage around rail stations, was modest, but I learned to read an execution plan and to treat performance as something to measure rather than guess.
At work, my responsibilities are ordinary but useful. I maintain SQL scripts that turn ticket data into daily ridership reports, answer questions from the scheduling team and check feeds when figures look odd. Last year I proposed, and then wrote, a small validation step that flags any vehicle whose upload is missing or unusually small before reports are generated. It is a Python script, not a sophisticated system, but it cut the number of corrected reports noticeably and taught me how much reliability depends on knowing what normal looks like. It also showed me the limits of my current knowledge: our setup is batch-based, and I would like to understand how stream-processing systems reason about event time, watermarks and late data, the ideas I first met in Martin Kleppmann's Designing Data-Intensive Applications, which I read on the train to work last winter.
Outside work, I run the Saturday morning junior chess club at my local library, which I joined as a player at eleven. Teaching eight-year-olds why a move is weak, without simply telling them the better one, has made me far more patient at explaining technical problems to colleagues who do not write code. I also play in a weekend league myself, with mixed results.
At postgraduate level I hope to study distributed storage, large-scale query processing and streaming in depth, and to gain experience with cloud infrastructure at a scale my current job cannot offer. I am comfortable with Python, SQL and Spark, and I have been working through linear algebra and probability revision independently, as I expect machine learning on large datasets to be part of the course. In the longer term I would like to work as a data engineer on public infrastructure such as transport, energy or health, where data quietly arriving late can affect real decisions. I am not looking for a change of direction so much as the depth to do my current kind of work much better, and I believe I would bring practical experience of messy, real data to discussions with my fellow students.