JoinGG

Profiling a Server That Feels Slow but Logs Nothing

A server that is late rather than failing leaves no trace in the logs. How to establish the fact, read a profile, and name the plugin responsible in a minute.

SweetMask · 7 min read

The server is up. The logs are clean. Nothing has crashed, nothing has errored, and players say it feels wrong.

This is the hardest category of server problem because every tool that reports failures reports nothing. The server is not failing; it is late, and lateness leaves no trace unless you go looking for it.

What "slow" means

A game server runs a loop on a fixed schedule. Every tick it works out where everything is, resolves what happened, and sends the result out. The budget is fixed: 50 ms at 20 ticks per second, about 15 ms at 64.

Late means the loop did not finish inside its budget. The server carries on, it does not error, it just falls behind, and players experience that as blocks reappearing, hits not registering, and movement that corrects itself half a second later.

So the first job is not to guess at causes. It is to establish whether the loop is late at all, because roughly a third of "the server is lagging" reports turn out to be the network between the server and the person complaining. Ping and latency covers that half, and it needs a completely different investigation.

Establish the fact first

Every game exposes some version of this number and the names differ.

Minecraft and its derivatives report ticks per second and milliseconds per tick. /tps on most plugin platforms. Twenty is healthy; anything under nineteen sustained is real.

Source engine servers report a server framerate through stats or the equivalent. Compare it against the configured tickrate.

FiveM reports per-resource timings in its own console, which is unusually useful because it attributes the cost directly.

Rust and Unity-based survival games expose a frame time in the server console.

If the number is healthy while people complain, stop here. The server is fine and the problem is elsewhere: their route, their machine, or a specific interaction that is not a general slowdown. Chasing server performance at this point is hours spent on the wrong thing.

Then find where the time goes

This is the step people skip, and it is the one that answers the question.

Use a profiler, not intuition. On plugin-based Minecraft, Spark. On Source, the engine's own profiling commands. On FiveM, the resource monitor. On a JVM generally, an async profiler attached to the process.

Run it for sixty seconds while the server is bad. Not while it is quiet, not for ten minutes — one minute under the conditions people are complaining about.

What comes back is a breakdown of where the tick spent its time, usually as a tree. Read it top down and look for the first thing that is not the engine itself. On a plugin server that is nearly always a plugin, named directly, with the method it was in.

This single step resolves the majority of slow servers, and it takes less time than reading an article about it.

The usual answers, in rough order

One plugin doing work on a hot event. Something running on every block place, every player move, or every tick. Frequently a plugin nobody has thought about in a year.

Entity count. Dropped items, mobs, minecarts, projectiles. Each is simulated every tick, and the distribution matters more than the total: one chunk with four thousand entities is a different problem from four thousand spread evenly. Why your Minecraft server lags covers the Minecraft-specific version in detail.

A database write in the wrong place. A plugin writing synchronously on an event that fires often. The signature is a stutter correlated with player actions rather than a steady sag, databases for game servers covers spotting it.

Chunk or region loading. Players spread across a map load more of it. Twenty players in twenty places is considerably more work than twenty in one town.

Garbage collection pauses, on JVM games. Visible as periodic freezes at regular intervals rather than continuous slowness. RAM, heap and Java flags covers why more memory usually makes this worse.

The machine. Last, and when it is the machine it is nearly always single-core speed rather than core count. More cores will not fix your server.

Reading the shape of the problem

The pattern tells you as much as the profiler does.

Constant slowness that scales with player count points at per-player work: entity tracking, view distance, or something running per connection.

Slowness that does not scale with players points at something running regardless; a farm, a clock, a scheduled task, an entity accumulation.

Periodic spikes at regular intervals point at a scheduled job: autosave, a backup, garbage collection, or a plugin timer.

Spikes correlated with specific actions point at an event handler. Somebody opens a shop, somebody teleports, somebody breaks a block with a plugin attached to it.

Gradual degradation over hours, fixed by a restart points at a leak or accumulation, and the restart is masking it rather than fixing it. Automating restarts covers when that trade is reasonable and when it is hiding something you should find.

Degradation over weeks points at data growth: a database table nobody prunes, a world that keeps expanding, logs filling a disk.

The measurements that cost nothing

Before any of the above, four commands that rule out whole categories in under a minute.

df -h for disk space. A full disk produces failures that look like anything else and this is free to check.

free -h for memory pressure.

htop for per-core load. One core near its ceiling with the rest idle is the normal and expected shape; several cores busy means something unusual is happening.

journalctl -k | grep -i "killed process" in case the kernel has been killing something.

Linux basics for game server owners covers these properly, and reading game server logs covers where the evidence usually is.

What to do with the answer

If it is a plugin: update it, configure it, or remove it. Most plugins doing expensive work have a setting that reduces the frequency, and the ones that do not usually have a maintained alternative.

If it is entities: set limits, clear what has accumulated, and find the farm producing them.

If it is chunks: reduce view and simulation distance, and set a world border.

If it is the database: move the write off the main thread if the plugin allows it, and prune the table.

If it is the hardware: buy clock speed. And before you do, check whether reducing the work is cheaper than increasing the capacity, because it usually is and it is free.

Reading a profile without being frightened by it

The output looks like a wall and it is a tree, and there are only three things to do with it.

Find the total, then find the largest child. The root is one tick or one second of ticks. Underneath it, the engine's own work and everything you added. Whatever holds the largest share below the root is where to look, and on a healthy server that is the engine.

Follow the largest child down until the name stops being the engine. Then you have a plugin, a resource, or a subsystem, with the method it was inside. That name is the answer, and it took four clicks.

Ignore everything under a few percent. A profile lists hundreds of entries and almost all of them are noise. A plugin at 1.2% is not your problem no matter how suspicious it looks.

The two traps are worth naming. Self time and total time are different: a plugin whose total is large because it called something expensive is not necessarily the plugin at fault. And a sample taken while the server is healthy tells you nothing about why it is unhealthy, which is why the timing of the capture matters more than its length.

When the profile says nothing is wrong

It happens, and it means something specific rather than that you failed.

If the tick is inside budget and players still complain, the loop is not the problem and no amount of further profiling will produce an answer. Move to the network path, and treat it as an unrelated investigation rather than a continuation of this one.

If the tick is late and the profile shows the engine itself dominating with nothing unusual underneath, the server is doing legitimate work beyond what the machine can do in the time available. That is a capacity answer: fewer players, less view distance, smaller simulation area, or a faster core.

If the tick is late intermittently and the profile never catches it, the event is shorter than your capture or triggered by something you are not reproducing. Log the slow ticks instead. Most platforms can record when a tick exceeded a threshold and what was running, and a week of that finds things a sixty-second sample never will.

If a restart fixes it every time, you have an accumulation rather than a workload, and the profile taken an hour after a restart is the wrong profile. Take one just before the next restart is due.

The habit worth forming

Profile once a quarter when nothing is wrong.

That sounds pointless and it is the single most useful thing on this page. Knowing what a healthy tick looks like on your server, which plugins cost what, what the entity count normally is, how much headroom exists — means that when something changes you can see it immediately rather than starting an investigation from nothing.

Most servers that become unplayable did so gradually, over months, in ways that were visible in a profile the whole time and that nobody was looking at.

Common questions

How long should a profiler run to capture server lag?
Run the profiler for sixty seconds while the server is actively experiencing problems. Capturing a profile while the server is quiet or letting it run for ten minutes will not give you the specific breakdown you need.
What should be done if lag spikes occur too briefly for a 60-second profiler capture?
Log slow ticks instead of relying on a short capture. Most platforms can record when a tick exceeds a specific threshold and identify what was running during that spike.
When server hardware causes late ticks, should you add more CPU cores?
No, you should buy faster single-core clock speed rather than increasing core count. Before purchasing hardware, check whether reducing workload like view distance or simulation area resolves the issue for free.
Tags
server adminperformancetick ratemaintenancetroubleshooting
Share

SweetMask

Published · 7 min read

All articles

Keep reading

Guides

Skyblock Servers, and Why the Format Refuses to Die

A map with a tree and forty blocks of dirt became one of the most played formats in multiplayer Minecraft. What the servers added, and which of the two Skyblocks you are joining.

· 7 min read