Who is doing what on your Splunk? User Activity Monitoring (UAM)¶
About this white paper
This white paper walks through User Activity Monitoring (UAM), the dedicated tenant type introduced in TrackMe 2.4.18 as a Beta: the inventory of the scheduled searches of every search head tier, the search activity of every account read from the audit trail, the resource usage of every search read from the introspection, the classification of accounts into people and services, and the findings the policies raise on top — with their investigation workflow and state-aware notifications.
It is a tutorial on a live deployment: a Splunk Cloud search head cluster monitoring itself, used by a handful of analysts and service accounts, to which we add a UAM tenant and follow it from its creation to its first decisions. No host, stack or person of the real deployment is named; the accounts are the personas of the tutorial.
UAM is a Beta: a complete tenant type whose policies, defaults and evidence shapes may still move between beta releases; feedback is welcome.
UAM is available on the Enterprise and Unlimited editions, their trials and the Developer license — not on the Foundation edition.
Product guide references: UAM — User Activity Monitoring, User Activity Monitoring — in depth, Creating a splk-uam tenant.
The problem: nobody sees who does what¶
A Splunk platform is shared by people and by machines, and after a few years nobody can say precisely who does what on it:
Personal schedules. An analyst creates a scheduled report to feed a dashboard, then changes team or leaves. The report keeps running under an account nobody owns any more — or stops the day the account is disabled, and a dashboard silently goes dark.
All-time searches. A search with no lower time bound reads every bucket of an index. Scheduled, it is a recurring cost; ad hoc, it is the
index=* earliest=0that pins a search head for an hour. Nobody sees it until the platform slows down.Runaway service accounts. A connector, a SOAR playbook or a reporting script polls the jobs endpoint every second, or dispatches searches in a loop after a deployment change. The audit trail records every call; nobody reads it.
Resource hogs. One account, human or not, consumes half of the search CPU of the platform with a handful of searches. The introspection knows, per process and per node; the information sits in
_introspectionfor two weeks and ages out.
Splunk keeps the evidence of all of this — the audit trail and the resource
introspection are on every deployment, Enterprise or Cloud — but reading it is a
project: the formats differ by version, an audit line is not a search, a search is not a
person, and the volumes are such that a naive search over _audit is itself a problem.
The requirement behind UAM was simple to state: read those sources once, cheaply, and
turn them into questions an administrator can act on — with the answer kept as evidence
in the index, not as a dashboard someone has to run.
Why the audit trail and the introspection¶
Three sources, all native, all forwarded to the indexing layer by every search head:
The audit trail (
index=_audit): every search Splunk grants and completes, with the account, the app, the search id, the time range searched, the run time, the events and results — and every REST access. It is authoritative and complete, but verbose: a single search leaves a dozen lines, and the REST polling of a client leaves hundreds.The resource introspection (
index=_introspection): per-process samples of every search process, on the search heads and the indexers — CPU, memory, elapsed time, and Splunk’s own type, mode and provenance of the search (scheduler,rest:jobs,UI:Dashboard:<name>…). It is the only place where the cost of a search is measured rather than estimated, and where a Splunk Web search is distinguishable from an API call.The knowledge objects of each search head tier: the saved searches with their owner, app, schedule and dispatch bounds, the users and the roles — read through the REST API, by exact path when a change signal says what changed, in full once a day.
UAM reads the three with bounded, incremental collectors — the audit trail from a durable checkpoint, never all time; the introspection window by window; the inventory by delta — and writes the result as compact, indexed evidence: one record per search and window, the REST access summarised per account, the resource usage per search, the inventory changes, and the findings’ transitions. The details are in the in-depth reference.
What UAM delivers¶
Capability |
What you get |
|---|---|
Inventory |
The scheduled searches, users and roles of every search head tier, with every change recorded: who created a schedule, when it changed, when it disappeared. |
Activity |
One audit record per search and window: account, app, search id, type, status, run time, events and results, the time range searched; the REST access of every account summarised; the lifecycle of every search correlated. |
Resources |
The CPU, indexer share, peak memory and elapsed time of every search, the environment’s search capacity measured hourly, and budgets derived from it. |
Classification |
People detected from Splunk Web activity and browser logins, services detected from sustained automation, system accounts seeded, rules and manual classes on top. |
Findings |
Eleven policies in three families — inventory, activity, resources — raising one finding per policy and subject, with evidence, a suggested remediation, an investigation status, an assignee, notes, a history and recurrence. |
Notifications |
State-aware emails on the transitions of the findings (new, reopened, escalated, resolved…), as a digest per run or as a thread per finding, with deep links. |
AI |
The AI Findings Advisor (interactive or scheduled), and the AI Assistant with a UAM context for tenant administrators. |
Cost |
Bounded collectors with validated budgets; compact evidence — one record per search and window, the REST access of an account summarised per batch. |
The environment of this tutorial¶
Our platform is a Splunk Cloud stack: a search head cluster used by a security team,
over the indexing layer of the stack. TrackMe runs on that search head cluster and the UAM
tenant monitors it — one search head tier, local, and the audit, introspection and
internal data the cluster forwards to the indexers. About twenty accounts use it:
ten people — analysts and engineers such as
alice.moreau,ben.okafor,tom.hendricksordavid.nakamura— who work in Splunk Web, own a few scheduled reports and run ad hoc investigations;six service accounts —
svc-siem-collector(a SIEM integration dispatching searches through the API),svc-soar-actions,svc-itsi-sync,svc-grafana-reader,svc-bi-exporterandsvc-backup-audit;jenkins-legacy, a retired CI account that still owns a schedule and never logs in;svc-trackme, the service account with administrative rights that owns the TrackMe tenants and runs their trackers;the built-in
adminand the Splunk context identities.
Around 135 saved searches exist on the tier, nearly all of them scheduled — the
platform’s own health alerts under nobody and the reports the teams created.
Step 1 — Create the UAM tenant¶
From the Virtual Tenants page, Actions → Create a tracking tenant opens the component chooser. The User Activity Monitoring card is the sixth, with its Beta badge:
The Basics step asks for the tenant identity — we call the tenant user-activity.
The Configure UAM step is where the shape of the tenant is decided:
Search head tiers — the local search head is selected; a remote account would add another tier of the same indexing layer (an Enterprise Security cluster, for instance). Every remote tier passes its connectivity test before the creation is allowed. The route account of the indexing-layer searches is derived:
localhere.Historical backfill — the range of audit and resource history collected before the live start, 30 days by default (none, or up to 90). The introspection retention is typically shorter than the audit one, so the resource history may end earlier.
Indexes & RBAC is the step every tenant ends with — with one recommendation specific
to UAM: give the tenant a dedicated service account with administrative rights as its
owner (svc-trackme here). The trackers run as the owner, the scheduled AI review runs
as the owner, and the owner must administer the tenant. Review shows the summary and
creates the tenant; the trackers are scheduled immediately and the card appears on the
Virtual Tenants page:
Hint
Automating it. The same creation is one REST call —
POST /trackme/v2/vtenants/admin/add_tenant with tenant_uam_enabled=true and a
uam_config object (tiers, scheduled, backfill_days) next to the usual
identity, indexes, owner and roles. See Creating a splk-uam tenant.
Step 2 — Day one: inventory, first activity, first findings¶
Within the quarter of an hour that follows, the workspace fills in:
the inventory tracker reads the saved-search changes, the users and the roles of the tier, and the listing tracker runs the first full listing — the Scheduled searches tab shows the schedules with their owner, the owner’s class, the cron expression and the enablement;
the activity tracker reads the audit trail from its checkpoint, five minutes at a time, and the Activity tab draws the timeline of the audit records — one per search and window — with the account, the app, the search type, the outcome and the run time;
the policies tracker evaluates the inventory policies on the schedules and the activity policies on the trailing fifteen minutes, and the first findings appear.
The backfill plans itself at its first hourly slot and walks the past thirty days backwards, newest first, in fifteen-minute buckets, for a little under an hour per slot; the tenant card shows its progress and the Collection health tab its edges — on our stack the audit history goes back the full thirty days, the resource history about two weeks, the introspection retention of the stack. The backfilled history matters beyond the timelines: it is where the people already working on the platform are detected (from their logins and their Splunk Web searches of the past weeks) and where the service accounts earn their seven days of automation-only activity, so the classification is usable on day one instead of a week later.
On day one, the findings are the inventory facts: the schedules dispatched all time
(all_time_schedule, nine of them — an old asset-table report among them), the
one real-time schedule (realtime_schedule), and no personal_schedule yet —
that policy needs the owners to be classified as human first, which the backfill and the
next step bring.
Step 3 — Reading a finding¶
A finding is one row per policy and subject. The Findings tab shows the backlog: the
tiles (open, high / critical, new in the range, assigned to me, resolved, dismissed), the
transitions timeline, and the table with the severity, the subject (a scheduled search or
an account), the policy, first and last seen, whether the latest run still observed it,
the status and the assignee. We open the high finding raised on alice.moreau — an
account that ran an all-time search, the one thing that should not happen on a shared
platform:
The Investigate view gives everything an administrator needs to decide:
the subject — here the account, its class and where the class comes from (human, detected), its environment; for a scheduled search it is the saved search, its app, its owner and the owner’s class, its tier;
the evidence — here the policy window, the number of all-time searches observed in it and the threshold in force (
max_all_time_searches: 0); for an all-time schedule it is the dispatch bounds the inventory read (dispatch.earliest_timeempty: Splunk’s default, no lower bound), with the honest caveat that inline SPL modifiers may still restrict the search;the policy’s suggested remediation — advisory: UAM never edits nor dispatches the observed searches;
Open record in search and Open events in search — the current KV record and every indexed transition of the finding;
the assignment and the history — every decision with who, when and the note;
the notes (TrackMe’s shared Notes, for context) and the AI Findings Advisor.
We assign the finding to tom.hendricks, who administers the platform, mark it
investigating with a note, and move on. Two decisions close a finding: Resolve
says the investigation concluded — if the condition is still there at a later run, the
policy reopens the finding (and the history says so: reopened by the policy);
Dismiss says the finding is accepted — permanently, or for a duration, after
which it comes back for another look. The bulk edit applies the same decisions to a
selection.
Step 4 — Who is a person, who is a service¶
The Users tab is the account-centric view: every account seen in the inventory or the activity, its class, its roles, its findings, its audit records and executions, its last activity and its apps. The tiles count the accounts, the inventoried users, the classified and the unclassified ones.
Most of the classification happened without us:
alice.moreau,ben.okaforandtom.hendrickscarry the human chip with a detected marker: each of them ran a search from Splunk Web (Splunk’sUI:*provenance in the resource records) or signed in to Splunk Web from a browser. One interactive signal is enough, and the class is sticky.svc-siem-collector,svc-soar-actionsand the othersvc-*accounts carry the service chip. Left to the detection, they would earn it after seven distinct UTC days of automation-only activity —rest:jobsand the scheduler — and nothing a person does, withdrawn by a single Splunk Web search; a naming convention gets there on day one (below).nobody,splunk-system-userandadminare system by default.
One account is left to us: svc-trackme, the owner of the tenant, which we classify
system together with the administrators’ accounts: they run the platform’s own
workload and must never appear in a finding. The service accounts are one rule —
Settings → Account classification takes ordered exact or wildcard rules (svc-* →
service), previews their effect on the known accounts before saving,
and lists the manual classes and the detections:
Precedence is always the same: a manual class, then the system defaults, then the first
matching rule, then the human detection, then the service detection. Two limits of the
beta are worth knowing: the automation scan reads the legacy key=value audit format
(a JSON-only _audit gives no automatic service detection), and the seven days must lie
inside the audit retention.
With the owners classified, the next policies run raises the personal_schedule
findings — at severity low, by design: three of Alice’s reports, two of Chen’s. They
are hygiene, not incidents; they sit at the bottom of the backlog until the team decides
which ones move to a service account.
Step 5 — Tuning the policies¶
The Policies tab is the catalogue: eleven policies in three families, six on by default. Each row shows the enablement, the severity, the thresholds in force, the exclusions and the counters of the last evaluation; the modal edits one policy and previews the draft against the current inventory before saving.
What we changed on our stack, and why:
Severity.
all_time_searchis high by default — an all-time search is the one thing that should not happen — and we keep it.all_time_schedulestays medium.Exclusions. The
trackmeandSplunk_SA_CIMapps are excluded from every policy by default (the monitoring and the data-model accelerations would otherwise flag themselves). Every policy also takes exact owner exceptions — where a legitimate all-time schedule, an acceleration owner for instance, would go.Scope and opt-ins. The audit policies evaluate human accounts by default and we keep it so. We switch on the four opt-in policies —
unclassified_schedule,dense_schedule,launch_burstat 50 distinct searches per window andcostly_searchat 60 seconds — because on a small platform their thresholds are easy to read; on a large one, leave them off until the baselines are known: a limit set today would be wrong tomorrow.Insight. The policy insight (from the Last evaluation cell) explains an evaluation: the population, why records were unknown, excluded or not applicable —
out_of_scope_service,excluded_app— the accounts behind the unknowns (classified in place), and a seven-day history of the counters. It is how we checked that the widened scope actually evaluatedsvc-siem-collector.
Every threshold is compared strictly, and every policy has its own exact exclusions; a stored configuration keeps its values across upgrades.
Step 6 — Resources: budgets from the capacity¶
The two resource policies — resource consumption (an account’s CPU seconds over the window against its allowance) and resource intensity (the count of individually heavy searches) — carry no absolute figure to maintain. The activity tracker measures the environment’s search capacity every hour from the introspection: the cores of the search hosts (the search heads and the indexers that ran a meaningful share of the search processes) and the smallest of their memories. From it, with auto-thresholds on:
the allowance of an account is 10 % of the search CPU capacity per fifteen-minute window;
a heavy search reaches 2 % of that capacity, or 10 % of the smallest search host’s memory;
an account whose share of the environment’s search CPU exceeds 40 % gets its finding raised one severity level — the share is evidence and a boost, never a finding on its own, so another account’s load never opens and closes findings.
The policies stay in a learning state until the tenant holds 24 hours of resource
history and a capacity record — the backfill counts — and the Policies tab says so. On our
stack the first resource finding came the next morning: svc-siem-collector over its
allowance, with a 55 % share of the environment’s CPU (the severity raised from medium to
high), its most expensive searches listed in the evidence — a playbook re-running a
two-week search every five minutes since a change. The account deep-dive (Step 4) shows the
CPU timeline by search type, the indexer share and the twenty most expensive searches;
Open in Search gives the rollup SPL. The third resource policy, the daily baseline
deviation, stays off in preview mode: it records what it would raise after seven fully
covered days, and we will look at it before enabling it.
Step 7 — Notifications¶
UAM notifies on the transitions of the findings — no per-policy alert, no wizard: one tenant-level setting in Settings → Notifications. We enable it with the new, reopened, escalated and resolved events (the defaults; dismissed and assigned stay off), a minimum severity of medium (the low personal-schedule findings stay out of the mailbox), every policy, our email delivery account, two recipients, and the digest mode:
In digest mode every policies run that produced a transition sends one email
listing them, grouped by event then severity, each with a deep link that opens the
Investigate view of the finding. The per-finding mode sends one email per transition,
threaded per finding — opened → escalated → assigned → resolved read as one
conversation in the mail client — the right choice for a small backlog with named owners.
Delivery is fail-open and retried until the transitions expire; a delivery failure never
stops the collection and is surfaced by the Guardian. The indexed trackme:uam:finding
events remain the third channel for any Splunk alert or SOAR playbook.
Operating the tenant¶
Collection health — every tracker with its last run explained, its diagnostics and its seven-day history; the activity checkpoint and its lag; the inventory datasets per tier; the audit windows of the last run; the measured capacity; the backfill with its progress, edges and controls (pause, resume, cancel, plan again, run now). A scheduled run never disables the page; Collect now dispatches every tracker on demand.
Recovery — a run killed mid-way leaves a lock; a lock older than five minutes whose search job has stopped on the same member is recovered automatically at the next run, or on demand with Recover stopped run.
Guardian — the
uam_collection_healthcheck (tenant scope, warning) surfaces a stale or failing tracker, a stalled checkpoint, an overdue or incomplete listing, a stopped backfill and a failed notification delivery on the Virtual Tenants page and in the Configuration Guardian, with the problems in its metadata; it clears itself once they are gone.Collection limits — every budget (activity windows and records, resources, inventory cadences and listing window, findings, backfill) is a tunable with a validated range, not a product cap; the defaults fit a busy cluster.
Cost and scale¶
Collection: bounded searches with validated budgets — the audit trail five minutes at a time from a durable checkpoint (never all time, a window split when it overflows), the introspection window by window, the inventory by delta with one full listing a day in an off-peak window, the backfill for a little under an hour per slot. A run that dies is replayed, never skipped; a degraded source is retried a bounded number of times, then recorded as a gap rather than stalling the cursor.
Evidence: one compact record per search and window — the REST access of an account summarised in one record per batch rather than hundreds of lines — one resource record per search and window, the inventory changes and one receipt per run. It counts against the Splunk license like every TrackMe summary event. Store SPL search stays off by default: the SPL adds about a third to a record, and it can carry literals and secrets.
Reads: the per-account and timeline views run through
tstatson indexed dimensions — O(accounts) rather than O(events) — and the account deep-dive fans its searches out concurrently with a bound, so a heavy account answers with what finished.State: three KV collections per tenant (settings, inventory, findings); the findings collection is capped (50,000 by default) and accelerated on its time fields.
What the beta does not do yet¶
Concurrent-use and off-hours policies — an account with N searches running at once, a human searching at 3 a.m. — need qualified start / end semantics of the search lifecycle and are not in the beta catalogue.
Notifications v2 — delivery to the assignee’s own address and a webhook channel remain future decisions; the beta delivers by email through the delivery accounts and by the indexed events.
Cross-environment identity linking — an account is a labelled source claim per tier; UAM does not link
alice.moreauon one deployment toamoreauon another, nor claim that a person sat at the keyboard.Automatic remediation — UAM never edits, disables or dispatches an observed saved search; every remediation is a suggestion for an administrator, and the AI Findings Advisor records decisions, not changes to the platform.
See also
UAM — User Activity Monitoring (Beta) and User Activity Monitoring — in depth — the product guide.
Creating a splk-uam tenant — the creation wizard and the REST flags.
Configuration Guardian — the Configuration Guardian.
The AI Advisors — the AI advisors, including the scheduled review of findings.
Using SLA alerting to build a 2-tier monitoring system — a two-tier alerting design for the entity components.