User Activity Monitoring (Beta)¶
Who is doing what on your Splunk: one tenant watches every account through the audit trail and the introspection, classifies accounts, raises findings from a policy catalogue and gives you the investigation workflow.
about 30 minutes · 31 slides. Each slide is followed by a description of what it shows — read top to bottom, this page is the walkthrough.
Overview¶
User Activity Monitoring (UAM) is the seventh tenant type of TrackMe, introduced in 2.4.18 as a Beta. One tenant watches every account of a deployment, people and service accounts alike, through the audit trail and the introspection, raises findings from a policy catalogue, and provides the investigation workflow. Everything shown here was captured on TrackMe 2.4.18 (Beta) on a Splunk Cloud search head cluster, starting from a new tenant called uam-cloud.
Four parts follow: creating the tenant, the first day of collection, reading and classifying what was found, then tuning the policies and operating the tenant. The accounts that appear, alice.moreau, ben.okafor, svc-siem-collector and others, are fictitious names of a small platform team.
The other half of the platform¶
Data monitoring sees the feeds. A Splunk platform is also a few hundred saved searches, twenty accounts and the automation behind them, and their behaviour has no owner: schedules nobody reviews, people and services that look alike in the audit lines, resource hogs that stay invisible until the search head suffers.
UAM is a dedicated tenant type. It reads the audit trail and the introspection, keeps an inventory of the scheduled searches, users and roles, records one compact evidence record per search and window, classifies every account as system, human or service, and evaluates eleven policies that raise findings with evidence, a suggested remediation, an assignee, a status and a history. It never edits or dispatches the searches it observes.
The collection and evaluation pipeline¶
Callouts 1 to 5 follow the pipeline. Sources: the audit trail (_audit), the introspection (resource usage, capacity), and the saved searches, users and roles of every tier. Trackers: inventory, listing, activity, policies and backfill, with bounded windows, durable checkpoints and one receipt per run. Evidence: trackme:uam:* records in the summary index, one per search and window. Classification: system, human or service, seeded, detected, then rules and manual classes. Findings: eleven policies in three families raise one finding per policy and subject.
UAM is not an entity component: its own wizard card, tabs and trackers, one tenant per deployment covering one indexing layer and one or more search head tiers. Investigation sits on top: assign, investigate, resolve, dismiss.
The four parts¶
Part 1, Setup, picks the UAM card in the tenant chooser, walks the Basics, Configure UAM, Indexes & RBAC and Review steps with a dedicated service-account owner, and ends on the tenant card. Day one follows the five trackers scheduled at creation, the collection health with its receipts and backfill, the tenant overview, and the first findings, which are inventory facts.
Part 2, Read & classify, covers the Findings tab and the Investigate view, the Users tab with its human, service and system classes, one account deep-dive, the Activity tab, and how the classification is decided. Parts 3 and 4 cover the policy catalogue and the insight, the classification rules and collection limits, the notifications, and collection health.
Choosing the UAM tenant type¶
The tenant chooser opens from Virtual Tenants, Actions, Create a tracking tenant. The User Activity Monitoring card (callout 1) is the sixth card, next to the five entity components; one tenant monitors one deployment. Its Beta badge (callout 2) marks a complete tenant type whose policies and evidence may still move between beta releases. Select (callout 3) opens the UAM flavour of the tenant wizard: Basics, Configure UAM, Indexes & RBAC, Review.
A UAM tenant carries no entities: no DSM, DHM or FLX component can be added to it, and it does not appear in the entity views. It has its own tabs: Findings, Users, Activity, Scheduled searches, Policies, Collection health and Settings.
The Basics step¶
The Basics step is the one every tenant type shares. The tenant id (callout 1), uam-cloud here, is an immutable identifier checked for availability as it is typed; it names the KV collections (settings, inventory and findings) and the trackers of the tenant. The alias and description (callout 2) are what people see on the Virtual Tenants card, “User activity monitoring” here. The Beta notice (callout 3) states what the beta does not do yet and that the schema may move between beta releases.
One deployment needs one UAM tenant. A second tenant would monitor a second indexing layer, another Splunk stack, not a second team of the same one: the RBAC of the findings lives inside the tenant.
The Configure UAM step¶
Configure UAM is the only step specific to UAM, and it decides the shape of the tenant. Search head tiers (callout 1): the local search head is the first tier, and any remote account adds another tier of the same indexing layer, an ES cluster for instance; a tier’s saved searches, users and roles are inventoried and its audit trail is read. Every remote tier must pass its connectivity test (callout 2). The route account (callout 3) runs the indexing-layer searches; derived, local here.
The historical backfill (callout 4), 30 days here, is the audit and resource history collected before the live start, up to 90 days. It makes the classification usable on day one instead of a week later.
Choosing a service-account owner¶
Indexes & RBAC is the step every tenant ends with. The indexes (callout 1) are the summary, audit, metric and notable indexes, where the defaults are fine; the trackme:uam:* evidence lands in the summary index. The owner (callout 2) is the recommendation specific to UAM: a dedicated service account with administrative rights, svc-trackme here, because the trackers and the scheduled AI review run as the owner. The roles (callout 3) are the admin, power and user roles of the tenant.
The owner must be able to read _audit and _introspection and to list the saved searches, users and roles of every tier; without those rights the inventory and the activity stay empty, and the Configuration Guardian reports it.
The Review step and creation¶
The Review step shows the whole shape before anything is written. Identity and tiers (callout 1): uam-cloud, one local tier, the route account and the 30-day backfill. What gets created (callout 2): three KV collections (settings, inventory, findings), the lookups, five trackers and the health tracker. Scheduled immediately (callout 3): the inventory, listing, activity and policies trackers start at their next slot, and the backfill plans itself at its first hourly slot.
Create (callout 4) completes in seconds; the collections, the lookups and the trackers exist immediately and are scheduled at once. The same creation is one REST call to vtenants/admin/add_tenant with tenant_uam_enabled and a uam_config object, for those who automate their tenants.
The UAM tenant card¶
Back on the Virtual Tenants page, a few seconds after creation, the tenant has its card, deliberately different from the entity tenants. The UAM card (callout 1) shows the alias, the identifier, the Beta badge and the tier list, with no entity KPIs and no donut. The open findings pill (callout 2) counts open findings by severity; it is empty right after creation, and the inventory facts appear within the quarter of an hour.
The backfill pill (callout 3) reads planned, running with its progress, or complete, as the history is walked backwards, newest first. The collection status (callout 4) stays green while the trackers run on time; the Guardian check uam_collection_health turns it amber when something stalls.
Trackers, receipts and backfill progress¶
Within the quarter of an hour that follows the creation, the Collection health tab shows the trackers (callout 1), inventory, listing, activity, policies and backfill, each with its last run explained and its seven-day history. Receipts (callout 2) are one trackme:uam:receipt per run: windows read, records written, gaps, duration. The activity checkpoint and lag (callout 3) show the audit trail read five minutes at a time from a durable checkpoint, never all time.
The backfill progress (callout 4) walks thirty days backwards in fifteen-minute buckets, under an hour per hourly slot, bounded by the retention of the sources; on this fresh stack it retired itself early. Collect now (callout 5) dispatches every tracker on demand.
The tenant overview on day one¶
The overview modal from the tenant card summarises the tenant. The KPI tiles (callout 1) count the accounts seen, the scheduled searches inventoried, the open findings by severity and the evidence records of the day. Transitions over time (callout 2) plot new, reopened, escalated, resolved and dismissed as a timeline of the backlog. By policy (callout 3) shows which policies raised what.
On day one the numbers are inventory facts: all_time_schedule, schedules dispatched all time, and realtime_schedule, a real-time schedule. There is no personal_schedule finding yet (callout 4): that policy needs the owners classified as human first, which the backfill and the classification bring. This is the starting point of every UAM tenant.
Part 2 — Read & classify¶
Part 2 reads the tenant. The Findings tab and the Investigate view give an administrator everything needed to decide on a finding, one row per policy and subject. The Users tab is where the classes settle: who is a person, who is a service, who is the platform itself, and how UAM decided it.
One account deep-dive and the Activity tab with the raw evidence complete the picture. Most of the classification happens without operator input.
The findings backlog¶
The Findings tab is the backlog. The tiles (callout 1) give the shape: open, high / critical, new in the range, assigned to me, resolved, dismissed. The transitions chart (callout 2) shows the backlog over time by event, one bar per policies run that changed something. The by severity and by policy charts (callout 3) show where the weight is. The filters (callout 4) cover status, severity, policy, subject type, assignee and “still observed by the latest run”.
The table (callout 5) lists severity, subject (a schedule or an account), policy, first and last seen, observed, status and assignee, with bulk edit. Each finding keeps its identity across runs: a persisting condition updates the same finding.
The Investigate view¶
The Investigate view is the whole case for one finding, here the high finding on alice.moreau, an account that ran an all-time search. The chips (callout 1) read high, open, account. The subject (callout 2) shows the account, its class and its source (human, detected). The evidence (callout 3) shows the policy window, the all-time searches observed and the threshold in force (max_all_time_searches 0). The suggested action (callout 4) is advisory only: UAM never edits nor dispatches the observed searches.
Open events, record and owner account (callout 5) link to every indexed transition, the KV record and the owner’s deep-dive. Assign, investigate, resolve and dismiss (callout 6) record each decision with who, when and a note.
Accounts and their classes¶
The Users tab is the account-centric view. The tiles (callout 1) count accounts, inventoried users, classified and unclassified. Human, detected (callout 2): alice.moreau, ben.okafor and tom.hendricks, from a Splunk Web search or a browser login. Service, detected (callout 3): svc-siem-collector and svc-soar-actions, from seven distinct days of automation only, or from the svc-* rule. System (callout 4): nobody, splunk-system-user and admin by default; svc-trackme and the administrators by decision.
The findings and evidence counts (callout 5) give findings, audit records, executions, last activity and apps per account. The class chip actions (callout 6) classify an account in place: svc-bi-exporter becomes service before it earns its seven days.
The account deep-dive¶
The account deep-dive fans its searches out concurrently with a bound, so a heavy account answers with what finished. Identity (callout 1): svc-bi-exporter, service by the svc-* rule, with its roles, its apps and its tier. Activity (callout 2) plots the executions over time by search type: API dispatches only, no Splunk Web, the profile of an export job. Resources (callout 3) show the CPU timeline by search type, the indexer share, the peak memory and the twenty most expensive searches.
Findings (callout 4) lists every finding where this account is the subject or the owner; as a service, it is judged against the capacity-derived allowance. Open in Search (callout 5) gives the rollup SPL of each panel.
Audit records as evidence¶
The Activity tab is the raw evidence: what the audit trail said, compacted. Records over time (callout 1) plot the audit records by account source: the UI, the REST API, the scheduler. By account (callout 2) shows who generates the records: the connectors dominate by count, the analysts by variety. The records table (callout 3) lists account, app, search id, type, status, run time, events and results, and the time range searched.
The evidence pills (callout 4) show one trackme:uam:observation per search and window, with the REST access of an account summarised per batch. The search lifecycle (callout 5) correlates dispatch, run and completion on the search id with the resource record of the introspection.
Classification precedence and beta limits¶
Human, detected: a search from Splunk Web (the UI:* provenance) or a browser login; one interactive signal is enough and the class is sticky. Service, detected: seven distinct UTC days of automation-only activity; one Splunk Web search withdraws it. Rules and manual classes: ordered exact or wildcard rules (svc-* to service) previewed on the known accounts; a manual class overrides everything. System: nobody, splunk-system-user and admin by default, plus the tenant owner and the administrators.
Precedence is always manual class, then system defaults, then the first matching rule, then human detection, then service detection. Two beta limits: the automation scan reads the legacy key=value _audit format, and the seven days must lie inside the audit retention.
Part 3 — Tune¶
Part 3 tunes the policy catalogue and the settings. Eleven policies in three families, six on by default, each with its severity, thresholds, exclusions and scope. The catalogue ships with defaults that fit a busy cluster, so tuning is a few decisions: which severity, which exclusions, which scope.
The policy insight explains an evaluation and is how a change is checked. The classification rules and the collection limits complete the settings.
The policy catalogue view¶
The Policies tab is the catalogue (callout 1): eleven policies in three families, inventory, activity and resources, one row each. The enablement (callout 2) shows six on by default, four opt-in, and the daily baseline deviation in preview mode. Severity (callout 3): all_time_search is high by default, all_time_schedule medium. Thresholds in force (callout 4) are compared strictly; the resource policies derive theirs from the measured capacity.
Last evaluation (callout 5) shows the counters of the last run and opens the insight. Edit and preview (callout 6) edits one policy and previews the draft against the current inventory. Here the four opt-in policies were switched on, launch_burst at 50 distinct searches and costly_search at 60 seconds.
Reading a policy insight¶
The policy insight opens from the Last evaluation cell and explains one evaluation. The verdict (callout 1) for all_time_search reads complete: every applicable account was judged, 3 evaluated, 3 matched, 2 outside the scope. Matched subjects (callout 2) are the accounts behind the findings, with the evidence that matched. Not applicable and excluded (callout 3) lists the system accounts the policy never judges, the Splunk context identity and the tenant owner, with the reason.
The unknowns (callout 4) are accounts without a class, classified in place from the insight. The seven-day history (callout 5) shows the counters run after run: proof that every classified person is judged, every run, and that the system accounts can never raise a finding.
The three policy families¶
Inventory policies read the schedules: personal_schedule (on, low), a scheduled search owned by a human account; all_time_schedule (on, medium), a schedule dispatched with no lower time bound; realtime_schedule (on, medium), a real-time schedule left running; unclassified_schedule (opt-in), a schedule whose owner has no class yet; dense_schedule (opt-in), a schedule denser than the threshold. Activity policies read the audit records: all_time_search (on, high), more than zero all-time searches; launch_burst (opt-in), more than N searches in a window; costly_search (opt-in), a search above a cost threshold.
Resource policies read the introspection against capacity-derived budgets: resource_consumption (on, medium), CPU seconds against an allowance; resource_intensity (on, medium), individually heavy searches; resource_deviation (preview), the daily baseline deviation after seven fully covered days.
Account classification settings¶
Settings, Account classification has three layers the operator controls and two UAM observes. Manual classes (callout 1) list every class set from a row menu, with who and when: svc-trackme and the administrators as system here. Rules (callout 2) are ordered exact or wildcard rules, svc-* to service, previewed on the known accounts before saving; the preview is the way to check that svc-* does not catch something it should not.
Detections (callout 3) show the human and service detections as they stand, with the signal that decided each one. The detection switches (callout 4) turn a detection off when a platform convention makes it wrong, a shared browser account for instance.
Tuning the collection budgets¶
Collection limits is where the cost of the tenant is decided. Activity budgets (callout 1) set the windows and records per run: the audit trail is read five minutes at a time, and a window is split when it overflows. Resources and inventory (callout 2) set the introspection window, the inventory cadences and the listing window, one full listing a day. Findings and backfill (callout 3) set the findings collection cap (50,000 by default), the backfill slot duration and the bucket size.
Store SPL search (callout 4) is off by default: the SPL adds about a third to a record and can carry literals and secrets. Every budget has a validated range (callout 5); these are tunables, not product caps.
Part 4 — Operate¶
Part 4 operates the tenant day to day. Notifications are state-aware emails on the transitions of the findings, one tenant-level setting rather than one alert per policy. The Collection health tab tells whether the trackers are on time and covers the recovery of a killed run, and the Guardian check reports on the Virtual Tenants page when they are not.
The Scheduled searches tab is the living inventory of the tier.
Configuring finding notifications¶
UAM notifies on the transitions of the findings with one tenant-level setting and no per-policy alert. Events (callout 1): new, reopened, escalated and resolved are the defaults; dismissed and assigned stay off. Minimum severity (callout 2) is medium, so the low personal-schedule findings stay out of the mailbox. Policies and recipients (callout 3): every policy, two recipients. The delivery account (callout 4) is an email delivery account of TrackMe; delivery is fail-open and retried until the transitions expire.
Digest or per-finding (callout 5) chooses the shape: one email per policies run that produced a transition, grouped by event then severity with deep links into the Investigate view, or one email per transition, threaded per finding.
Receipts, diagnostics and recovery¶
Operating the tenant is mostly reading the Collection health tab. The tracker run view (callout 1) shows every tracker with its last run explained and its seven-day history, the measured capacity, and the backfill with its edges and controls. Receipts (callout 2) record what each run read and wrote: a run that dies is replayed, never skipped; a degraded source is retried, then recorded as a gap. Diagnostics (callout 3) explain a slow or partial run.
Recover stopped run (callout 4): a lock older than five minutes whose job has stopped is recovered at the next run or on demand. The Guardian check uam_collection_health (callout 5) surfaces a stale tracker, a stalled checkpoint or a failed delivery, and clears itself.
The living inventory of schedules¶
The Scheduled searches tab is the living inventory of the tier: 134 schedules here (callout 1), read by delta with one full listing a day. Owner and owner class (callout 2) show who owns a schedule and whether that is a person, a service or the platform; the inventory policies read this column. Schedule and dispatch (callout 3) show the cron expression, the enablement, and the earliest and latest time of the dispatch. App and tier (callout 4) show where a schedule lives; excluded apps are listed, just not evaluated.
Every change is recorded (callout 5) as trackme:uam:inventory events: who created a schedule, when it changed, when it disappeared. The tab answers how many schedules exist and who owns them.
Evidence footprint and beta scope¶
The cost of a UAM tenant is the evidence it indexes, designed to be small. Compact: one record per search and window rather than one line per audit event, REST access summarised per account and batch, one receipt per run. Bounded: five-minute audit windows from a durable checkpoint, introspection window by window, one full listing a day, backfill under an hour per slot. Licence: the trackme:uam:* evidence counts against the Splunk licence like every TrackMe summary event, Store SPL search off by default.
Beta statement: a complete, supported tenant type; the service detection needs the legacy key=value _audit format and seven days of audit retention. Not in the beta: concurrent-use and off-hours policies, notifications v2, automatic remediation.
Where we are¶
A UAM tenant was created in six screens, with one local tier, a 30-day backfill and a dedicated service-account owner, and five trackers scheduled at once. Twenty accounts were classified: people detected from Splunk Web, service accounts by one wildcard rule, the platform’s own accounts as system. The four opt-in policies were switched on. The tenant operates with digest notifications above medium, self-recovering runs and the uam_collection_health Guardian check.
Three things make it work. Backfill first: thirty days of history make the people and the services known on day one. A catalogue, not a wizard: eleven policies with defaults that fit a busy cluster. Evidence you can afford to keep, and a Guardian that says when collection stalls.