Guides
Evaluation criteria at a hackathon: how to judge fairly and without scandal
How to assemble a jury, create a clear scale, and not fall out with participants at the award ceremony.

The loudest conflicts at hackathons happen not because of prizes, but because of a sense of injustice. A team invested two sleepless nights, lost, and didn't understand why. Transparent evaluation criteria and a properly assembled jury are not bureaucracy, but insurance against a scandal that will ruin the event's reputation.
What makes up a fair evaluation
There is no single scale, but a working set of criteria usually looks like this. It is important to decide in advance how much weight each point carries and tell the participants before the start, not after.
- Usefulness and relevance. Does the project solve a real problem.
- Technical implementation. What exactly works, rather than just being drawn on slides.
- Completeness. Is there a working prototype, rather than just an idea in words.
- Presentation. Was the team able to clearly demonstrate the result.
- Originality. How unexpected the approach is.
The sum of the weights must equal one hundred percent, and the weights should reflect the goal of the hackathon. If you gathered product teams, usefulness weighs more than code cleanliness. If it is a purely engineering challenge, it's the other way around.
Whom to invite to the jury
A jury consisting only of top managers who do not write code evaluates beautiful slides and the speaker's charisma, not the project. A jury of only seniors digs deep into the architecture and fails to see the usefulness. Balance is more important than star status: a product person, an engineer, and a representative of the problem domain cover different angles.
Be sure to address conflicts of interest. If a mentor from a sponsor company judges a team they guided all night, people will notice. Either they do not evaluate this team, or they do not evaluate at all.
How to judge when everyone is coding with AI
Previously, the jury partly evaluated the volume of code written. Today, this is pointless: over a weekend with Cursor or Claude, one person can build what used to take a whole team. Shift the weight from "how much was written" to "what and why." Does it work on a live example, is a real problem solved, is there a thought behind the project rather than just a three-screen wrapper over someone else's model. Ask the team what they understand about their own code; this quickly separates those who built it from those who just pressed a button. We have a separate breakdown on evaluating AI projects.
How the pitch session goes
Give each team the same amount of time and keep to it strictly. Three minutes for the demo and two for questions is fair to everyone. A team you accidentally gave seven minutes to gets an unfair advantage, and it shows.
- Distribute the scale and criteria to the jury in advance.
- Show a timer that is the same for everyone.
- Collect scores for each criterion, not just a general impression.
- Calculate the result transparently, and discuss controversial cases before the announcement, not after.
Manually calculating scores in a general spreadsheet is exactly where resentment and mistakes are born. On Stavleak, the jury has a separate login and an evaluation form based on your criteria. Each judge scores by points, the platform sums them up automatically, and the ranking is published immediately after evaluation. It suggests the winners based on the total score, leaving you only to confirm and issue certificates.
After the announcement
Give teams feedback, especially the losers. A single phrase from the jury about what was missing turns the bitterness of defeat into a clear lesson. People forgive losing, but they do not forgive silence and the feeling that everything was decided behind closed doors. You can set up transparent evaluation for your criteria on the Stavleak demo.