Perspective · Security operations architecture

The lake is not the SOC

Purpose-built SecOps platforms with a unified security data lake vs the generic organisation-wide data lake, and why the real question is who owns the security decision loop.

Walk into almost any large enterprise today and you will hear the same ambition: one data platform for everything. Finance, customer, product, operations and, increasingly, security telemetry, all landing in a single organisation-wide lakehouse. Add a layer of AI agents on top, and the thought follows naturally: why do we need a separate security platform at all?

It's a fair question, and it deserves a better answer than "because SIEMs have always existed." Modern enterprise data platforms are genuinely excellent at storage, large-scale analytics, machine learning, natural-language querying and, yes, threat hunting. Anyone who dismisses them as "just a data lake" hasn't looked recently.

But after nearly two decades of building and modernising security operations, from running a 24x7 SOC floor to architecting AI-driven platforms across South Asia, ASEAN and ANZ, I've come to believe the debate is framed wrongly.

The question isn't "enterprise lake or security lake?" It's which data needs to sit inside a real-time decision-and-response loop, and which can stay where it already creates value.

Storing is not operating

Security operations isn't a query. It's a chain that has to run continuously, correctly and at machine speed:

telemetry → security context → detection → correlation → investigation → decision → action

A generic data platform is superb at the first link and increasingly capable at the next few. The differentiation, and the risk, sits at the bottom of the chain. That's where a purpose-built SecOps platform with a unified security data lake earns its place.

Four things you inherit or have to build

A security-native data modelNormalisation is not a one-off ETL job. Vendors change log schemas without notice; open standards like OCSF help, but someone still owns parsing, mapping and the regression tests that prove detections still fire.
Entity resolution over timeA username, an email, a SID, a laptop, an IP and two cloud identities are one person. Stitching them continuously, and then stitching events into a single attack story, is what turns seven queries into one incident.
Maintained detection contentLLMs help write detections; they don't remove detection engineering. Thresholds, baselines, false-positive tuning and MITRE ATT&CK coverage need an owner as adversaries and telemetry evolve.
Safe, governed actionIsolating an endpoint or disabling an identity needs authorisation, approvals, blast-radius limits, rollback and an audit trail. Analytics does not equal enforcement.

None of these is impossible on a general-purpose platform. The honest point is that every one of them becomes your engineering team's product to build, run and support at 3 a.m.

AI raises the stakes, not the floor

It is now remarkably easy to build an agent that answers "is this user compromised?" by querying identity, endpoint, DNS and cloud logs and writing a polished summary. That demo is no longer the hard part.

The hard part is trust. A SOC agent can't be judged on whether its answer sounds right. It must be tested for hallucinated conclusions, missed evidence, wrong causal links, privilege boundaries and prompt injection. That includes instructions planted inside the very telemetry it reads, a risk the OWASP Top 10 for LLM Applications ranks first. Frontier models are becoming interchangeable components; what differentiates the outcome is the security context, tools, permissions and decision framework wrapped around the model.

The moment an agent moves from recommending a containment action to executing it, agentic AI stops being an assistant and becomes part of your security control plane. That is exactly where purpose-built matters most.

Don't choose. Tier.

The most mature architectures I see don't pick a side. They keep the enterprise data strategy and add a security operating layer, deciding deliberately where each class of data earns its keep:

Data classWhere it belongs
Hot, operational security telemetryThe security decision plane, for real-time detection, correlation and response
Historical and forensic dataLower-cost enterprise or object storage, searchable when needed
Business context (asset criticality, customer tier, fraud signals)Stays in the enterprise estate; accessed or selectively enriched
Low-value, high-volume logsNot duplicated by default

The enterprise lake has real advantages here: business context that can dramatically sharpen security reasoning, mature data-science teams, and retention economics. The goal is to use them, not compete with them.

Five questions worth asking

  1. Are we replacing log storage and search, or the whole detect–investigate–respond operating model?
  2. Who continuously owns detection content, normalisation and regression testing?
  3. How are identities, devices and cloud resources resolved into one attack context?
  4. How will we evaluate agents for hallucination, injection and model changes before they act autonomously?
  5. Once something is judged malicious, which control actually stops it?

Answer those honestly, include the engineering and 24x7 support in the total cost, and the architecture usually designs itself.

An enterprise lake can become part of your SOC. The question is who owns the security decision loop.

Views expressed here are my own and do not necessarily represent those of my employer.

Rethinking your SOC architecture?

I'm always happy to compare notes on SecOps data strategy and agentic security operations.

Start a conversationConnect on LinkedIn ↗