top of page
Search

RunReveal vs. Splunk vs. Microsoft Sentinel: One Real Detection, Three Languages

1 day ago
11 min read



Two things also changed under my feet while I was writing. ClickHouse bought RunReveal (announced 1 September 2026), and Microsoft is in the middle of moving Sentinel out of the Azure portal and into the Defender portal.


A comparison that ignores either of those would be out of date the day it went up.


Where my knowledge comes from: RunReveal: hands-on, everything in this article was run against a real workspace this week. Splunk and Sentinel: their current official documentation and pricing pages, checked in October 2026.

What Changed in 2026

RunReveal is now part of ClickHouse

ClickHouse announced it had acquired RunReveal on 1 September 2026. That isn't as odd, RunReveal was always built on top of ClickHouse, with every log landing in one big columnar logs table. ClickHouse's own announcement says existing contracts and support stay the same, and that RunReveal remains available through its bring-your-own-database model, where your security data sits in a ClickHouse cluster you control

The practical effect for anyone evaluating it today: the pricing page no longer lists the tiers I recorded back in September. It now says RunReveal has joined ClickHouse, that a joint offering is being worked on, and to contact sales. I'll come back to what that means for cost later on.



Sentinel is moving house

Microsoft originally planned to retire the Sentinel experience in the Azure portal on 1 July 2026. That date has since moved to 31 March 2027. After that, you manage Sentinel in the Microsoft Defender portal. New customers have been sent to the Defender portal by default since July 2025, and Microsoft's docs say Sentinel works there even without Defender XDR or an E5 licence.

The bigger change is underneath.

Sentinel now has two storage tiers:

the analytics tier (the classic Log Analytics data that analytics rules, hunting and incidents run on) and a data lake tier in open Parquet format that keeps data for up to 12 years. Microsoft's docs say analytics-tier data is mirrored into the lake, so you keep one copy. You query the lake with KQL jobs or Jupyter notebooks, and promote data back into the analytics tier when you need detections to see it.


Splunk is a Cisco product with three ways to pay

Cisco completed its Splunk acquisition in March 2024, and the pricing story has spread out since then. Splunk's own pricing FAQ now lists three models for Enterprise Security: ingest-based, workload-based and activity-based.

Splunk Cloud Platform offers all three;

self-managed Splunk Enterprise offers ingest or workload. On the query side, Splunk Cloud's Federated Search for Amazon S3 lets you search S3 data without indexing it first. It uses SPL2, not classic SPL, and Splunk positions it for archival or low-value data rather than real-time detection.


The Short Version

  • Query language:

    RunReveal: ClickHouse SQL, plus Sigma rules

    Splunk: SPL (and SPL2 for federated search)

    Microsoft Sentinel: KQL


  • Storage engine:

    RunReveal: ClickHouse, columnar, one wide logs table

    Splunk: Splunk indexers and buckets (SmartStore in the cloud)

    Microsoft Sentinel: Log Analytics (analytics tier) plus a Parquet data lake tier


  • Detections:

    RunReveal: Scheduled SQL or Sigma, defined as text from the start

    Splunk: Correlation searches in Enterprise Security

    Microsoft Sentinel: Analytics rules (scheduled and near-real-time KQL)


  • Allowlisting:

    RunReveal: SQL WHERE clause or an enrichment tag

    Splunk: Lookup tables

    Microsoft Sentinel: Watchlists


  • How you pay:

    RunReveal: Was storage-based ("no ingest fees"); now "contact us" after the ClickHouse deal

    Splunk: Ingest, workload or activity-based

    Microsoft Sentinel: Analytics tier per GB (pay-as-you-go or commitment tiers) plus data lake charges


  • Where it fits best:

    RunReveal: SQL-fluent teams, cloud and SaaS-heavy log sources

    Splunk: Large enterprises, deep app ecosystem, existing SPL skills

    Microsoft Sentinel: Microsoft 365 and Azure-heavy shops


Architecture: Where Your Logs Actually Live


RunReveal

One table does most of the work. Every source lands in logs, with a handful of normalized columns (sourceType, eventName, actor, srcIP, eventTime, receivedAt) and the original event kept in rawLog. You query the normalized columns when they're enough and dig into rawLog with ClickHouse JSON functions when they aren't. Written article showed the one rule that matters most for speed: filter on receivedAt, because that's what the table is ordered by.


Every query in this article does that, and none of them took longer than about 0.8 seconds over a week of data.

Splunk

Splunk parses events into indexes made of time-bucketed files, and extracts most fields at search time. That schema-on-read approach is why Splunk copes with almost any messy log you throw at it, and why the add-on (TA) for each source matters so much: the TA is what turns raw text into the field names your searches use.

Splunk Cloud's SmartStore keeps the bulk of indexed data in object storage with a local cache, which changes the cost profile but not how you write searches.

Microsoft Sentinel

Sentinel sits on Log Analytics workspaces. Each connector writes to its own table, so instead of one big logs table you get dozens of purpose-built ones. That's great for well-supported Microsoft sources and more awkward for third-party ones, where the table you end up with depends on which version of the connector you installed.

Google Workspace is a good example.

The older Azure Functions connector writes to custom tables such as GWorkspace_ReportsAPI_token_CL. The newer schema uses a GoogleWorkspaceReports table, which Microsoft's table reference marks as a Basic log table that supports lake-only ingestion. Same data, different tables, different query.


One Real Detection, Three Languages

Here's the problem I picked, because it's a real one: a third-party OAuth app gets granted the full Gmail scope, https://mail.google.com/. That scope means read, send and delete for the whole mailbox. Consent phishing campaigns go after exactly this, so most Google Workspace detection packs include some version of the rule.


First, what the data actually looks like

Google records each grant as an authorize event under the token application. The interesting parts (app name, client ID, scopes) aren't top-level fields. They sit inside an array of name/value pairs in events[0].parameters, which every platform has to unpack somehow.

"events": [{
  "name": "authorize",
  "parameters": [
    {"name": "client_id",   "value": "1173955..."},
    {"name": "app_name",    "value": "..."},
    {"name": "client_type", "value": "WEB"},
    {"name": "scope_data",  "multiMessageValue": [...]},
    {"name": "scope",       "multiValue": ["https://www.googleapis.com/auth/..."]}
  ]
}]

RunReveal (ClickHouse SQL),

My first attempt picked parameters out by position (the first one is client_id, the second app_name, the fifth scope). It worked, and it's a trap: nothing in Google's API promises that order, and if a parameter is ever added or dropped the rule silently starts reading the wrong field.

This version finds each parameter by its name instead:

WITH JSONExtractArrayRaw(JSONExtractRaw(rawLog,'events'),1,'parameters') AS params,
     arrayFirst(p -> JSONExtractString(p,'name') = 'app_name', params) AS p_app,
     arrayFirst(p -> JSONExtractString(p,'name') = 'scope',    params) AS p_scope
SELECT JSONExtractString(p_app,'value') AS app_name,
       count()                  AS grants,
       uniqExact(actor['email']) AS users,
       uniqExact(srcIP)          AS ips
FROM logs
WHERE receivedAt > now() - INTERVAL 7 DAY
  AND sourceType = 'gsuite' AND eventName = 'authorize'
  AND has(JSONExtract(p_scope,'multiValue','Array(String)'), 'https://mail.google.com/')
GROUP BY app_name
ORDER BY grants DESC

Run against workspace, it came back in under half a second with exactly one app: our email-security vendor's integration, holding the full Gmail scope for 8 users, with 432 grants from 274 different IP addresses over the week. Both versions, by position and by name, returned the same 432, which is how I checked that the name-based one wasn't missing anything.

Read that result cold and it looks alarming:

a mailbox-wide scope, used from hundreds of IPs. It's actually expected. Cloud email-security products scan mail through the Gmail API from a pool of cloud addresses. That's exactly why this rule needs an allowlist before it goes anywhere near a notification channel, which is the next section.



Splunk (SPL), untested translation

In Splunk, the work happens in field extraction. The Splunk Add-on for Google Workspace collects the Reports API activity, including the token endpoint. How the parameters array comes out as fields depends on the add-on version, so treat the field names below as placeholders and check them against your own events first:

index=gws "events{}.name"="authorize" "https://mail.google.com/"
| spath path="events{}.parameters{}" output=params
| mvexpand params
| eval pname=spath(params,"name"), pval=spath(params,"value")
| eval app_name=if(pname="app_name", pval, null())
| stats values(app_name) AS app_name, values(actor.email) AS user, values(ipAddress) AS ip BY id.uniqueQualifier
| stats count AS grants, dc(user) AS users, dc(ip) AS ips BY app_name
| lookup oauth_app_allowlist app_name OUTPUT approved
| where isnull(approved)

It's noticeably more work than the SQL, and most of that is unpacking the array (spath, then mvexpand, then pulling name and value back out). In fairness, a well-built add-on usually flattens this for you at search time, and then the rule shrinks to a couple of lines. That's the

Splunk trade in a nutshell: a lot depends on the add-on, and when the add-on is good, writing searches is pleasant.

Microsoft Sentinel (KQL), untested translation

KQL handles the nested JSON cleanly with mv-apply, which loops over the parameters array per event. Again, the column names depend on which Google Workspace connector populated your table, so adjust them before using this. I've written it against the raw JSON so it doesn't depend on any one connector's flattening:

let allowlist = _GetWatchlist("OAuthAppAllowlist") | project SearchKey;
GoogleWorkspaceReports      // or GWorkspace_ReportsAPI_token_CL on the older connector
| where TimeGenerated > ago(7d)
| extend raw = todynamic(RawEvent)   // RawEvent: whichever column holds the original JSON
| extend ev = raw.events[0], Actor = tostring(raw.actor.email), Ip = tostring(raw.ipAddress)
| where tostring(ev.name) == "authorize"
| mv-apply p = ev.parameters on (
    summarize AppName = take_anyif(tostring(p.value), tostring(p.name) == "app_name"),
              Scopes  = take_anyif(p.multiValue,      tostring(p.name) == "scope")
  )
| where set_has_element(Scopes, "https://mail.google.com/")
| where AppName !in (allowlist)
| summarize Grants = count(), Users = dcount(Actor), Ips = dcount(Ip) by AppName

The watchlist is Sentinel's answer to Splunk's lookup table: a CSV you upload and reference by name, which is a nice way to let someone outside the detection team keep the list of approved apps up to date.


A Sentinel catch worth knowing: Microsoft's table reference marks GoogleWorkspaceReports as a Basic log table with lake-only ingestion supported. Check which tier your data is landing in before you build a scheduled analytics rule on it. Microsoft's own data lake docs say you can't run analytics rules directly against data lake tables; the data has to be in the analytics tier, or promoted there by a job.


The Noise Problem Is the Same Everywhere

Before I narrowed the rule to the Gmail scope, I looked at every OAuth grant in the workspace. The numbers are the most useful thing in this article, because they'd hit you on any of the three platforms:


  • Total authorize events (7 days): 24,364, about 145 an hour

  • Distinct OAuth apps behind them: 2

  • Grants from the SIEM's own Google collector: 23,703 (97%), one client ID, 1 user, 7 cloud IP addresses, present in all 168 hours

  • Grants from the email-security integration: 661, 8 users, 295 IP addresses

  • Scopes the collector asks for: admin.reports.audit.readonly and apps.alerts


97% of the OAuth activity in this tenant is the log collector getting a fresh token so it can read the audit log. It requests admin-reports scopes, every hour, from a rotating set of cloud IPs. Write the common rule "alert on OAuth grants with admin scopes" and the first thing it catches is your own SIEM, about 140 times an hour.


Splunk and Sentinel are no different here:

their Google Workspace connectors poll the same Reports API with the same kind of service account, so they should generate the same events about themselves.

This is also the second time this series has found the collector being the noisiest thing in the data.



How each platform lets you exclude it

  • RunReveal: put the exclusion straight into the SQL (AND client_id NOT IN (...)), or tag the collector's events with an enrichment and filter on the tag. Remember from Part 4 that enrichment only tags events from the moment it's activated; it doesn't go back and tag old ones.

  • Splunk: a lookup table of approved apps or client IDs, joined in the search. Easy to share across many correlation searches.

  • Sentinel: a watchlist, as in the KQL above, or an automation rule that closes matching incidents. The watchlist is better, because the noise never becomes an incident at all.


Query Languages, Honestly

SQL is RunReveal's biggest advantage, and also where it costs you something. If your team already writes SQL, the ramp-up time is close to zero, and ClickHouse's JSON and array functions are fast and capable. But security-specific idioms you'd get for free elsewhere (a time-chart command, transaction grouping, a built-in lookup operator) you write yourself as SQL.


SPL is a pipeline language built for exactly this job. It takes real time to learn, and it has quirks (multivalue fields and mvexpand bite everyone at some point), but people who are fluent in it are very fast, and there's a huge amount of shared content written in it. SPL2 brings it closer to SQL, but in practice you'll still be writing classic SPL for most detections.


KQL feels the nicest of the three to read. It pipes like SPL, it's typed, and it treats dynamic JSON as a first-class type. If your team already uses KQL for Defender advanced hunting or Azure Monitor, that skill carries straight over.


Detection as Code

In RunReveal, a detection is text from the start: SQL or Sigma, created and updated over the API. Previous Article showed a detail that matters if you sync detections from git: a Sigma rule is stored as raw YAML in settings.rule, and its query field is empty. A naive export that only reads query will back up every Sigma rule as a blank.

I also hit a real gap in the API while testing: deleting a detection through the API didn't work, which is why every test query in this article ran ad hoc instead of being saved.

In Splunk, Enterprise Security correlation searches are saved searches underneath, so they can live in an app's configuration files and be versioned like any other config. Splunk also publishes an open library of security detections that many teams use as a starting point.


Sentinel analytics rules export as ARM templates, and Sentinel's Repositories feature can deploy rules straight from a GitHub or Azure DevOps repo. That's a solid story if you already run infrastructure as code in Azure.


Ingestion: Breadth vs. What You Measure

RunReveal's catalogue is around 120 connectors. That's fine for cloud, SaaS and identity sources, and thinner than Splunk's add-on ecosystem, which covers almost anything that writes a log.

Sentinel's catalogue is big too, and unbeatable for Microsoft's own sources, many of which come with free ingestion.


What I'd look at more than the connector count is lag and duplication, because those affect detections directly.

RunReveal's Google Workspace polling lag at between 17 and 155 seconds (median 83), not the round 60 seconds you might assume. Every API-polling connector, on every platform, has some version of this delay. Measure it on whichever platform you pick, and set your rule look-back windows to fit.


AI and Automation

RunReveal built AI into the workflow early:

AI-assisted query writing in Explorer, triage on investigations, and scheduled Agents that run a prompt on a cron schedule and report back.


There's a catch I ran into first-hand: those features need an AI provider configured on the workspace, either your own API key or a cloud role RunReveal can assume. With nothing configured, the agents can't run. That's a reasonable design for data control, but it means "AI-native" comes with setup and a bill of its own


Splunk offers an AI Assistant for writing and explaining SPL.
Sentinel connects to Security Copilot, and the Defender portal move pulls Sentinel incidents into the same queue as Defender XDR.

Microsoft's data lake docs also lean on notebooks and AI as reasons to keep years of data. The ClickHouse announcement says RunReveal's agentic investigation work is meant to shape how ClickHouse supports agentic analytics, so expect that side to get more attention, not less.


Running It Day to Day

This is the part feature lists skip.

  • With RunReveal there's very little to run: no indexers, no forwarder fleet, and most of my time went into writing queries rather than looking after the platform.


The flip side is that it's a younger product. I hit undocumented behaviour in nearly every part of this series (string-versus-boolean enabled fields, a delete endpoint that didn't delete, enrichment that doesn't backfill), and you'll find these by testing, not by reading the docs.

  • Self-managed Splunk is the opposite: a lot of infrastructure, but very few surprises left in it after this many years. Splunk Cloud takes most of the infrastructure away.

  • Sentinel needs no infrastructure, but you do need to understand Azure: workspaces, data collection rules, RBAC, and now the Defender portal on top of all of that.


So Who Should Pick What

  • Pick Splunk if you need the widest ingestion coverage, you already have SPL skills or content, or you run a lot of on-prem and unusual log sources. Go into the pricing conversation knowing all three models, and model your search load as well as your ingest.

  • Pick Microsoft Sentinel if Microsoft 365, Entra ID and Azure are most of your estate. The free Microsoft data sources and the Defender integration are hard to beat. Plan the Defender portal move now rather than in March 2027, and decide early which tables go to the analytics tier and which go to the lake.

  • Pick RunReveal if your team thinks in SQL, your sources are mostly cloud and SaaS, and you want detections as code without running infrastructure. Before you sign, get the post-acquisition pricing in writing, and ask where the hosted product is going, since ClickHouse's announcement emphasises the bring-your-own-database route.


Whichever you pick, do what I did here before trusting any rule: run it against a week of your own data, group by whoever is generating the events, and see how much of it is your own tooling talking to itself. In this tenant it was 97%.

-------------------------------------------------------------Dean-------------------------------------

 
 
 

Comments


Ready to discuss:

- Schedule a call for a consultation

- Message me via "Let's Chat" for quick questions

​

Let's connect!

Subscribe to our newsletter

Connect With Me:

  • LinkedIn
  • Medium

© 2023 by Cyberengage. All rights reserved.

bottom of page