The situation
The platform was already in production and used daily by around 2,000 students. It had grown quickly, and it showed: one of the most-used pages took 10 to 15 seconds to load, search results didn’t behave the way users expected, and notifications weren’t always reaching the right people. At the same time, new feature requests kept arriving.
Rewriting the system wasn’t an option. Students were using it every day.
What I did
- Found and fixed the slowest page first. I traced where the load time was going and optimized that page, since it affected the most users.
- Improved search behaviour so results matched what students were actually looking for.
- Fixed the notification flows end to end, from the backend trigger to what the user sees.
- Shipped UI improvements and new features alongside the fixes, working within the existing Flask/FastAPI architecture.
- Diagnosed production bugs as they were reported.
- Coordinated user acceptance testing (UAT) so each release was checked by real users before going live.
The result
- The key page went from 10–15 seconds to about 5 seconds.
- Releases kept shipping on a steady schedule, with no rewrite and no downtime for students.
What this means for you
If your application is slow or unreliable but people depend on it daily, a rewrite is rarely the answer. Measuring where the time goes and fixing the highest-impact problems first usually gets results in weeks, not months.