Executive Presence on Camera: Zoom and Video Call English for Saudi Real Estate Professionals
Quick Answer
Most high-stakes conversations with foreign investors now happen on video before, and sometimes instead of, in person. Yet video strips away the physical cues, the handshake, the room, the shared coffee, that usually carry warmth and authority, which means your English delivery has to work harder to do that job alone. This guide covers a simple framing and lighting method, how to look present through a camera lens rather than a screen, pacing adjustments for latency, how to narrate a screen share without losing energy, how to read disengagement on video, and the exact English phrases for handling technical problems without losing composure or credibility.
Introduction: The Venue Has Changed, the Stakes Haven't
A growing share of the conversations that decide whether a foreign investor moves forward on a Saudi property now happen entirely over video. Roadshow follow-ups, first introductions to a family office based in London or Singapore, weekly updates to an institutional buyer's team, even parts of negotiation, increasingly happen through a screen rather than across a table. The commercial stakes of these conversations are identical to an in-person meeting. The tools available to make a strong impression are not.
In person, a significant amount of trust and authority is carried by things that require no English at all: a firm handshake, calm posture, the shared experience of walking through a property together. On video, almost none of that is available. What is left is your face, your voice, your framing, and your English delivery, which means each of those has to work considerably harder to carry the same impression. This is a distinct, learnable skill, not simply "the same meeting, but online," and treating it as identical to in-person presence is where many otherwise strong professionals lose an edge they do not realize they are losing.
Why Camera Presence Is a Different Skill From In-Person Presence
Three structural differences separate video from in-person communication, and each has a direct English-delivery consequence.
Body language compresses. On video, only your head, shoulders, and hands (if visible) carry non-verbal signal, compared to your entire posture and presence in person. This means facial expression and vocal tone have to carry proportionally more of the warmth and confidence that used to be spread across your whole body.
Latency changes turn-taking. Even a small audio delay disrupts the natural rhythm of conversation, causing people to talk over each other or leave awkward gaps. Natural in-person turn-taking cues, a breath, a slight lean-in, often do not translate cleanly over video.
The lens is not the screen. The single most disorienting adjustment for most professionals: the camera lens and the screen showing the other person's face are in two different physical locations. Looking at the person you are speaking with, which feels natural, actually looks like you are looking down or away on their end. True eye contact on video means looking at the lens, which feels unnatural and slightly performative at first, precisely because it is not where the visual information you want is located.
Understanding these three mechanics is what turns "I'm just not as good over video" into a specific, trainable set of adjustments.
The Frame-Light-Eye Method
A simple three-part check before any important video call resolves the majority of on-camera presence issues.
Frame. Position the camera at or slightly above eye level, never looking up at you from below. Leave a small amount of headroom above your head without excessive empty space, roughly the same framing a professional headshot would use. Choose a clean, uncluttered background, or a subtle blur, that reads as a serious professional environment rather than a casual one.
Light. Face your primary light source rather than sitting with it behind you. A window behind you creates a silhouette effect that makes you difficult to read and can look evasive on screen, regardless of what you are actually saying. A simple desk lamp or ring light facing you, even a chair turned to face a window instead of away from it, solves this in seconds.
Eye. Look directly at the camera lens, not at the image of the other person's face, whenever you are delivering a key point, an answer to a direct question, or your headline number. This is the hardest habit to build and the highest-leverage one: eye-lens contact reads as directness and confidence on the other end, in a way that looking at the screen, however natural it feels to you, simply does not replicate. A practical technique: place a small visual reminder, a sticky dot, directly beside your camera lens during practice sessions until looking there becomes automatic.
Voice and Pacing for Video: Adjusting for Latency
Even a fraction of a second of audio delay is enough to disrupt natural conversational rhythm, and the adjustment that compensates for it is almost entirely about pacing. Speak slightly slower than feels natural in person, leave a beat longer before responding to make sure the other person has actually finished, and avoid jumping in immediately after what sounds like the end of their sentence, since a small pause that feels awkward in person often simply means the other side's audio is still catching up.
Because body language reads smaller on screen, vocal energy has to compensate. A flat, monotone delivery that might be balanced out by confident posture and gesture in person reads as considerably less engaged on video. Deliberately widening your vocal range, more variation between emphasis and normal delivery than feels necessary, closes a gap that viewers otherwise perceive as low energy or disinterest, even when the underlying content is identical.
Narrating a Screen Share Without Losing Energy
The moment a screen share begins, many professionals unconsciously shift into a quieter, more monotone "reading mode," precisely because the other person can no longer see their face and the instinct to perform confidently seems to switch off. This is exactly the wrong moment to lose energy, since the content being shared, a pricing table, a yield model, a location map, is frequently the most commercially important part of the entire call.
The fix is active narration: describe what you are doing as you do it, not just what is on the screen. "I'm scrolling down to the yield sensitivity table now, this is the section I mentioned earlier" keeps the audio engaging even though the visual has changed, and it also helps anyone whose screen briefly lags or freezes stay oriented. Maintain the same vocal energy and pacing adjustments from the previous section throughout the share, and consciously return to camera-visible delivery, ending the share and re-establishing eye-lens contact, at the moment you deliver your most important concluding point.
Reading the Virtual Room
Disengagement looks different on video than in person, and missing the signals costs you the chance to re-engage before a call is effectively over even though it has not ended. Watch for a camera that stays off longer than the opening minutes would explain, noticeably delayed responses to direct questions, eyes that repeatedly shift away from the camera area in a pattern that suggests a second screen or device, and a general flattening of verbal engagement, shorter answers, fewer follow-up questions, than earlier in the call.
The English recovery move is the same one that works in person, adapted for video: pause, check in directly, and invite honesty. "Let me pause here for a second, does this match what you were expecting, or would it help to go in a different direction?" This does two things at once: it interrupts a drifting call naturally, without confrontation, and it signals that you are reading the room, which itself rebuilds engagement in a way that simply talking louder or faster never does.
Handling Technical Problems Gracefully in English
Technical disruptions are inevitable on video calls, and how they are handled in English says as much about your professionalism as the content of the call itself. Composure under a frozen screen or a dropped connection reads as executive presence; visible frustration reads as the opposite, regardless of whose connection is actually at fault.
Useful, calm phrasing for common situations: for audio delay or overlap, "Sorry, go ahead, I think we spoke at the same time"; for a frozen or frozen-looking video, "You've frozen on my end, I'll give it a moment before we continue"; for someone clearly unable to hear, "I don't think you can hear me, I'll try reconnecting and message you if it doesn't resolve in a minute"; for your own connection issue, "Apologies, my connection dropped for a moment, could you repeat the last point?" None of these phrases apologize excessively or draw unnecessary attention to the disruption; each simply names what happened and states the next step calmly, which is exactly what keeps a technical hiccup from becoming a credibility issue.
The First and Last Sixty Seconds
Video calls remove two moments that normally do significant relationship-building work in person: the walk to the meeting room and the walk out afterward, both of which offer natural, low-stakes small talk that builds warmth before and after the substance of a meeting. On video, that warm-up and cool-down have to be deliberately built into the first and last sixty seconds instead, or they simply disappear.
Open with genuine, brief warmth before moving to the agenda, a real question about their location, their travel, or a light follow-up from a previous conversation, not a rushed "shall we get started." Close with equal deliberateness: summarize the agreed next step clearly, thank them specifically for their time, and avoid the awkward, drawn-out video-call goodbye by ending decisively once the substance is complete. Both moments take under a minute and meaningfully shape how the entire call is remembered.
Common Mistakes
| Mistake | Better Approach |
|---|---|
| Looking at the screen instead of the camera lens | Look at the lens deliberately when delivering key points |
| Sitting with a window or light source behind you | Face your light source; reposition before the call, not during it |
| Going quiet and monotone during a screen share | Actively narrate what you're doing, maintain vocal energy throughout |
| Jumping in the instant someone appears to finish speaking | Leave a slightly longer pause than feels natural before responding |
| Ending a call abruptly with no summary or warmth | Deliberately build a brief warm close into the final sixty seconds |
Pre-Call Checklist
- Check your frame, light, and camera height before every important call, not just the first time
- Confirm your background is clean, professional, and free of distracting movement or clutter
- Do one full run-through of your key numbers or points looking directly at the lens
- Prepare your calm, standard phrases for a dropped connection or audio delay in advance
- Plan a brief, genuine opening moment and a clear, warm closing line before the call starts
- Test your screen-share content and internet connection a few minutes before the call, not during it
Frequently Asked Questions
1. Why does looking at the camera lens feel so unnatural? Because the lens and the image of the other person's face are in different physical locations on your screen, so true eye contact requires looking away from the face you are actually watching. It becomes automatic with deliberate practice.
2. Do I need to look at the lens the entire call? No, natural glancing at the screen is fine and expected throughout most of the conversation; the technique matters most specifically when delivering a key point, answering a direct question, or making your most important statement.
3. How much should I slow down for video calls? Enough to comfortably avoid interrupting through latency, typically a noticeably longer pause before responding than feels necessary in person. It should feel slightly deliberate to you while sounding completely natural on the other end.
4. What's the biggest energy mistake during a screen share? Going quiet and monotone the moment the share begins, precisely when the content being shown is often the most commercially important part of the call. Active narration keeps energy consistent throughout.
5. How do I know if someone has genuinely disengaged versus just being naturally quiet? Look for a pattern, not a single moment: sustained camera-off time, consistently delayed responses, and shorter answers together, rather than one quiet stretch, which can simply be normal listening.
6. What should I say when my own connection drops? Something calm and brief once you reconnect: "Apologies, my connection dropped for a moment, could you repeat the last point?" Avoid over-apologizing, which draws more attention to the disruption than the disruption itself.
7. Does this apply the same way to a one-on-one call and a large group investor call? The core techniques apply to both, though reading disengagement is harder in a large group, since any individual camera or response pattern is less visible; deliberate, direct check-ins matter even more in group settings.
8. Should I always keep my camera on for the full call? In most professional real estate and investor contexts, yes, camera-on is the expectation and default; going camera-off for an extended stretch without explanation can itself read as a disengagement signal to the other side.
9. How do I build a stronger closing without it feeling scripted? Anchor it to something specific from that call, a particular question they asked or point they raised, rather than a generic closing line; specificity is what keeps a warm close from sounding rehearsed.
10. What's the fastest way to improve video presence before a real investor call? Recording yourself in a practice run and reviewing it critically, ideally with a coach, closes the gap between how you think you come across on camera and how you actually do faster than any other single exercise.
Summary
Video removes almost every physical cue that normally carries trust and authority in person, which means your framing, lighting, eye-lens contact, pacing, and vocal energy have to do that work instead. The Frame-Light-Eye Method, deliberate pacing for latency, active screen-share narration, reading disengagement early, and composed English for technical hiccups together close the gap between how strong a professional you actually are and how strong you come across on a screen.
Continue the series: previous, How to Explain Contracts in Plain English: A Guide to SPAs, MOUs, and Key Clauses for Saudi Real Estate Professionals; next, Investor Relations English: Writing Earnings Updates and ESG Reports for Saudi Real Estate Professionals.
About the author. Bilel Shelbi is the Founder of BEOS (Business English Of Substance), a Canadian native English speaker of Algerian origin, fluent in Arabic and French, with more than a decade of corporate language coaching experience and a top-2% international ranking. BEOS delivers confidential 1-on-1 deal-communication coaching for GCC real estate professionals, 100% remote via Zoom.
Book Your Confidential Communication Performance Assessment
A private 45-minute evaluation of exactly where your English breaks under deal pressure, and a personalized roadmap.
Meet the coach →