Zumigo Blog

Beyond Doc-V: Why Identity Verification Needs Contextual, Phone-Based Signals

Gartner’s newly released 2026 Magic Quadrant for Identity Verification confirms something the market has known for a while: document verification and biometric liveness are now table stakes. Investment in presentation attack detection (PAD), injection attack detection (IAD), and GenAI-content detection is critical, and it’s paying off — the AI-driven threats targeting these checks are very real. Doc-v has never been more accurate at answering one question: is this a real document, held by a real, live human, in front of this camera, right now?

But that question, however well it’s answered, is not the same as the question an enterprise actually needs answered: is this the right human, using their own identity, in a session the enterprise can trust?

Gartner’s own report draws this distinction directly. In the “Context” section, the analysts write that PAD, IAD, and inspection for GenAI-created images “are not enough in the current aggressive threat landscape,” and that the leading vendors “have invested in a range of additional risk signals, such as those focused on device, phone number, email, behavior and location.” The report goes further, noting that not every impersonation attack even involves a deepfake. Some are simpler: a real person, live and present, using someone else’s real, stolen identity. A document check and a liveness check can both pass cleanly in that scenario, because neither one asks whose identity the live human is actually holding. That’s the gap, and it’s exactly the gap that contextual, phone-based signals are built to close.

Why a passed doc-v check can still be a compromised session

Consider what doc-v and liveness actually verify: the document is genuine, the face matches the document, and a live human is present. All three can be true and the transaction can still be fraudulent:

  • The device or SIM was compromised before the check ever ran. A SIM swap that happened hours earlier hands an attacker the phone number tied to the account. From that point forward, every OTP, every SMS verification link, every “we texted you a code” step is being delivered straight to the attacker — while the person taking the selfie is, genuinely, alive and holding a valid ID.
  • Call forwarding is silently active. The legitimate customer’s number still exists, but calls and texts are being redirected. Doc-v never looks at this. It has no visibility into carrier-level session state at all.
  • The session itself is spoofed. A deepfake video, a virtual camera, or an injected image can be defeated by strong liveness detection — but an attacker running that attack from an unfamiliar device, in an unfamiliar location, with no history tied to the real customer’s profile, produces a device fingerprint that doesn’t match anything doc-v is designed to check.

In each case, the document-and-face check can pass cleanly. The risk is sitting one layer below it, in signals doc-v was never built to see.

Check the phone before you even send the OTP

Most verification flows still start from the same assumption: send a one-time passcode or a verification link to the phone number on file, then trust whatever comes back. That assumption is exactly what a SIM swap or call-forwarding attack is designed to exploit — the OTP goes out, and it goes straight to the attacker, who now clears every downstream check including doc-v.

The fix isn’t a better OTP. It’s not sending one blind in the first place. Before an SMS link or OTP goes out, the phone number itself should be evaluated:

  • Has this SIM changed recently? Real-time SIM swap detection can catch a takeover precursor in the window before the OTP is even generated — before the attacker gets the chance to intercept anything.
  • Is call forwarding active on the line? A number can be technically “in service” for the real customer while every call and text is quietly redirected elsewhere. This is invisible to a document check and invisible to liveness, but it’s directly visible to carrier-level signals.
  • Does the line itself carry risk? Account tenure, porting history, and line type (a disposable prepaid line activated last week looks nothing like a number with years of stable carrier history) all say something about the identity behind the number before a single message is sent.

This reframes the order of operations: phone risk assessment happens first, as a gate, not as an afterthought bolted on after doc-v has already passed.

Mobile network authentication: verifying the phone before any actions are allowed

Once the line itself checks out, there’s a second, complementary layer: mobile network authentication. Silent Network Authentication (SNA) confirms, directly with the carrier and without any customer action, that the SIM in the device making the request actually belongs to the registered subscriber on the account — no code to type, no link to tap. SIM-based Authentication (SBA) extends that same carrier-level confirmation to networks or devices where a silent check isn’t available, asking only for the user to acknowledge their phone number rather than requiring a full OTP round-trip. Between the two, the phone gets verified on nearly every network and device — fully silently where possible, with a single lightweight step where it isn’t.

That distinction matters. An OTP only proves that whoever holds the phone at this moment can receive a message. Mobile network authentication — whether silent via SNA or acknowledged via SBA — proves that the phone itself, at the carrier level, is genuinely tied to the identity being claimed. It’s a fundamentally different, harder-to-forge signal, and it’s one no synthetic identity or session hijack can produce — because there’s no real subscriber relationship behind it to confirm.

Device fingerprint: does this session look like this customer?

The third layer is the device and session context itself. Device fingerprinting, IP intelligence, and location consistency answer a question doc-v never asks: does this session look like the customer’s established pattern of behavior, or does it look like someone else operating from an unfamiliar device, network, or geography?

An attacker can present a stolen or synthetic identity that sails through document and liveness checks. What they can’t do is make their device, IP, and location quietly match months or years of the real customer’s established profile. That mismatch is often the clearest signal available — and it’s sitting in a completely different data layer than anything a doc-v vendor evaluates.

The layered answer: doc-v plus context, not doc-v alone

None of this is an argument against doc-v. Document authenticity and biometric liveness remain a necessary foundation — they’re the only layer that directly answers “is there a real, live human here, and does the document check out?” That foundation keeps getting pushed forward against AI-fabricated documents and deepfakes, and it needs to.

The point is that foundation alone leaves a gap that sophisticated fraud — and, per Gartner’s own reporting, plenty of decidedly low-tech impersonation — walks straight through. The layered answer looks like this:

  1. Gate the phone number before anything is sent to it — SIM swap and call-forwarding checks, line intelligence, and tenure signals, run before an OTP or SMS link ever goes out.
  2. Verify the phone itself, at the network level — SNA confirms the carrier-level relationship silently; SBA extends that same confirmation, with a lightweight user acknowledgment, wherever a silent check isn’t available.
  3. Verify the session context — device fingerprint, IP intelligence, and behavioral consistency confirm the session matches the customer’s established pattern.
  4. Then, and only then, layer in doc-v and liveness as the check that the physical document and the live human presenting it are genuine.

Enterprises that treat a passed doc-v check as the finish line are trusting a single signal — one that AI has gotten increasingly good at spoofing, and one that isn’t sufficient on its own even before AI enters the picture. The organizations getting this right are the ones treating identity verification as a stack of independent, corroborating signals — document, biometric, and contextual — where no single layer has to be perfect, because the others are watching for exactly what it misses.

Learn more about our solutions to combat to the unique threats posed by AI to enterprise identity here.

 

Madhu Vudali is VP, Product Management at Zumigo. Comments or questions? Connect on LinkedIn: @madhuvudali