A five-minute screen recording can easily be larger than a five-minute filmed video shot on a phone, which seems backwards until you consider what a video codec is actually good at compressing. Codecs rely heavily on finding redundancy between frames and smooth, gradual color transitions, both things screen content is unusually bad at providing.
Why screens are a worst case for video compression
Natural video, a face, a landscape, tends to have soft gradients and continuous motion that compress efficiently, since neighboring pixels are usually similar and change gradually between frames. A screen recording is nearly the opposite: sharp, high-contrast text edges, flat colored UI elements, and a cursor that can jump instantly across the frame, all of which the codec has to represent precisely to keep text legible, since even minor blurring makes small text unreadable. Scrolling is a particularly bad case, it changes almost every pixel in the frame at once, defeating the inter-frame compression that saves the most space in typical video.
How to shrink a screen recording without losing readability
The most effective approach is capping frame rate rather than slashing bitrate: screen content rarely needs 60fps, since most of what's being demonstrated (clicking, typing, navigating) reads perfectly well at 15-24fps, and the Video FPS Changer can reduce this after the fact for a recording that was captured at an unnecessarily high frame rate. Resolution matters too, but cropping too aggressively risks making text illegible, so it's worth testing with the Video Bitrate & Resolution Changer rather than guessing. If the recording was made with the Screen Recorder, starting with a reasonable capture resolution in the first place avoids needing to recover lost detail during compression later.

