Optimising the Next.js application
6 min read
October 06, 2026
All of of a sudden, we had the opportunity to conduct a state-level online examination for orthopaedics students across Tamil Nadu. We were told to expect more than 1000 concurrent users. At that point, the product was still in development and testing was happening alongside it. Our tech stack was Next.js and PostgreSQL, hosted on a KVM2 VPS (2 vCPU, 8GB RAM). Everything had been put together to work well enough to meet the demo criteria. It was nowhere near ready for 1000 concurrent users.
Our error margin was zero because it was a state-level examination. Any discrepancy, bug or glitch during the exam could have damaged our reputation permanently. The experience had to be as seamless as possible from start to finish.
I used k6 to run the stress tests. First, I had to turn "1000 concurrent users" into a number the server actually has to handle. Students open three pages during the exam, and on average a student stays on a page for about 40–45 seconds to read everything. So 1000 users generate roughly 1000 ÷ 42 ≈ 24 page requests per second. I set 30 pages/s as the practical target to keep a safety buffer. This matters because if you run k6 without simulating that reading time, every virtual user requests pages back to back, and the load jumps to over 400 pages/s. That doesn't reflect real student behaviour, and no server in our budget could handle it anyway. I also set 50 pages/s as a stretch target for the worst moment: the start of the exam, when everyone logs in at once.
I had no idea whether our VPS could handle these numbers. Before starting, I configured the server to protect the other containers running on it from failing during the test. I allocated 4GB of RAM to this project's container so it couldn't exceed that limit.
I began by ramping up towards 200 concurrent users at 10 pages/s, planning to hold for three minutes. The pages performed poorly. The app started failing before the ramp even reached 50 users, and it couldn't hold them for one minute. I SSHed into the server to monitor the stats and found that more than 80% of the container's RAM was being consumed by the Next.js process rendering pages. That's significant for a project and use case like this. I realised the poor results were down to how the project was built, not where it was hosted. In the first optimisation cycle, I went through every file to find the bottlenecks.
When I checked each page, rendering was taking more than 170ms of CPU time per request. That's definitely not normal. The CPU limit on our plan was fixed and wouldn't change unless we upgraded. I decided to optimise the app rather than upgrade the server, because no matter which server you pick to meet the requirement for the moment, the tech debt in the project doesn't go away. If anything, it hurts performance in the long run and costs more than it should.
The idea was to make the server do as little work as possible per request. Pages and sections that were the same for every student were moved to static rendering, so the server didn't rebuild them on every visit. I also cut down unnecessary client components, which reduced the JavaScript bundle and the rendering work done on the server for each page's initial HTML. After this round, I tested again. The app could now hold more than 200 users at 10 pages/s. I realised every test result depended heavily on how little time the server spent on each request. In other words, my focus had to be on optimising the pages.
In the next run, I increased the load to 400 concurrent users and reduced the sleep time to reach 15 pages/s. Throughout the tests, the database itself never maxed out, since we were using a serverless Postgres Neon provider that scales on demand. But it was hosted in Southeast Asia - Singapore, while our server runs in Mumbai, so every query paid a cross-region round trip. After this test, I decided to cache database query results and remove redundant queries. I also reduced the size of each page. Both changes cut the time the server spent per request and pushed me closer to my goal.
As testing went further, finding the remaining bottlenecks took longer and longer. Along the way, I found a few security gaps, unoptimised assets and redundant code, and fixed them. Honestly, optimising an application to get the outcome you need is a different ball game. After all these changes, the app could sustain 16 pages/s for 8 minutes.
With all these improvements, 1000 users were finally within reach. By then, the CPU time to render a page was under 50ms, down from over 170ms. That was a far better position than where I started. The app could now sustain more than 27 pages/s, which is above the realistic average load of about 24 pages/s for 1000 students. For the login spike at the start of the exam, which could go beyond that. I'm leaving out one variable here: the ramp-up pattern I used to increase users over time, because I changed it many times across tests.
I got these results through many iterations and experimented with a lot of variables. In early 1000-user runs, before I set a memory limit for the Node process itself, the server started swapping memory in and out. At other times, the CPU sat 10–25% idle while handling 1000 users. A single Node process mostly runs JavaScript on one core, so I assumed the second core was going unused. I set up clustering with PM2 so that two Node instances could serve pages at the same time. It didn't work out the way I expected. Two Next.js instances meant two copies of the app in memory inside the same 4GB limit, which brought the swapping back, and both instances were still competing for the same two vCPUs. Looking back, the idle CPU was more likely time spent waiting on database round trips than an unused core, and caching was the real fix for that, not more processes.
Throughout this process, I learnt that knowing which knob to turn and when, based on monitoring, is super important. It also taught me to build features with the server's limits in mind. During development, we shouldn't only focus on passing test cases. We should also think about how the server will handle each operation.