User Activity Monitoring — in depth¶
Tip
This is the in-depth reference for User Activity Monitoring (splk-uam, Beta in 2.4.18). For the readable overview of what UAM is and when to use it, start with UAM — User Activity Monitoring; for a guided, end-to-end walkthrough on a live deployment, read the User Activity Monitoring white paper. This page covers the tenant model, the search head tiers, the trackers and their receipts, the indexed evidence, identity and classification, the policy catalogue, the findings lifecycle, notifications, the operations, the REST endpoints, the tenant options and troubleshooting.
The tenant model¶
A UAM tenant is a dedicated tenant type: it carries
tenant_uam_enabledand no entity component. Creation refuses any mix with an entity component (uam_tenant_must_be_dedicated), the Manage components menu is disabled on its card, replica trackers are refused, and the add / remove component endpoints answeruam_tenant_is_dedicated.There are no entities: no decision maker, no impact scoring, no green / orange / red state, no stateful alerting, no priority / SLA / logical group. The unit of work is the finding; the alerting object is the finding’s transition.
Evidence lives in indexes, state in the KV store. Three per-tenant collections:
kv_trackme_uam_settings_tenant_<tid>(the configuration, the checkpoints, the latest receipt and the lock of every job, the capacity, the detection records, the notification outbox),kv_trackme_uam_inventory_tenant_<tid>(one row per saved search / user / role, per tier) andkv_trackme_uam_findings_tenant_<tid>(one row per finding). Each has an administrative lookuptrackme_uam_<kind>_tenant_<tid>.Access is tenant-admin only: every UAM read and write requires
trackmeadminoperationsand membership in the tenant’s admin roles (oradmin_all_objects). UAM evidence is audit data — the AI Assistant context of a UAM tenant is restricted the same way. Reads work on a read-only license; writes require a valid license beyond the Foundation edition.The tenant card shows the open findings by severity (the Tenant Card Detail Level preference applies: compact mode shows open findings and the combined high / critical count), the tiers and the collection status; a double-click opens the read-only overview (collection strip, seven-day transitions timeline, tiles, findings table).
The tenant identifier of a UAM tenant is short — lowercase letters, digits and hyphens, up to 20 characters — because it keys the per-tenant collections and reports.
Search head tiers¶
One indexing layer, N search head tiers:
Data |
Lives on |
Collected |
|---|---|---|
Saved searches, users, roles |
each search head tier |
inventory and listing, once per tier |
|
the indexing layer (every tier forwards there) |
activity, resources, capacity, once per tenant, through the route |
tiers— 1 to 10 unique names,localand / or remote account names (see Remote Splunk deployments). Every selected account must feed the same indexing layer; that invariant is the administrator’s — one remote account per tier of that layer.account— the route of the indexing-layer searches:localwhen selected, else the first tier; derived when a submission carriestierswithoutaccount.Typical shapes: on Splunk Cloud, TrackMe on the ad-hoc search head (
local) plus a remote account for the Enterprise Security search head; on Splunk Enterprise, any number of remote accounts over the same indexers; for central monitoring, remote accounts only.A user present on several tiers is one account (roles united, tiers listed); a saved search belongs to its tier. Each tier has its own inventory report, lock and receipt, so a slow tier never starves another and a skipped tier is visible on its own.
A search head that keeps its internal indexes local (no forwarding of
_audit/_introspection) hides its searches, logins and resource usage from the tenant — the forwarding prerequisite applies to every signal alike.
The trackers¶
Every job is a scheduled report running | trackmeuamtracker tenant_id=<tid> job=<job>
as the tenant owner, with its own lock (never stolen; a lock whose search job is
provably stopped is recovered after five minutes), its receipt (the outcome of the
latest run, kept in the settings collection and indexed as trackme:uam:receipt) and its
entry in the tenant’s operational status. Collect now in the workspace dispatches
job=all (every tier’s inventory, then activity, then policies) as the caller.
Report |
Job key |
Schedule |
|---|---|---|
|
|
|
|
|
daily in the off-peak window (01:00 + 5 h, search head local time) |
|
|
|
|
|
|
|
|
|
|
(not a UAM job) |
|
Inventory and listing¶
The inventory of a tier is the saved searches (with their owner, app, cron schedule, dispatch bounds, enablement and ACL namespace — SPL and alert actions are not read), the users and the roles, paged through TrackMe’s bounded REST reader (the local session or the remote account). A full listing costs the tier’s size whatever the filter, and an exact-path read is about ten times cheaper, so the collection is split:
inventory:<tier>(every 15 minutes) reads the changes reported by the tier’s own write log since the last run, by exact path (up toinventory_delta_max_objects— more escalate to a listing), and the users and roles everyinventory_identity_minutes. While the write log cannot be read (_internalnot searchable by the tier’s account, a member not forwarding it) the listing cadence falls back toinventory_fallback_reconcile_hours.listing:<tier>runs the full listing everyinventory_reconcile_hoursinside the off-peak window, with its own time budget (inventory_listing_max_seconds); the first listing runs within the hour of the creation. List now dispatches one.
Each record is projected to a KV row under the tier’s environment and diffed against the
previous row: a new row is added; changed field states changed (with
changed_fields); a classification change alone reclassified; a row seen again
after removal reappeared; a row absent from a complete scan is marked
present=false with removed_at — never deleted. A partial or unavailable dataset
upserts what was read and marks nothing removed. Change events are written only after the
batch save succeeded, so a failed write never publishes a change twice. The receipt
carries, per dataset, the status, source rows, pages, added / changed / removed / unchanged
counts and diagnostics.
Activity¶
The activity job reads the tenant’s _audit trail from a durable ingestion-time
checkpoint (cursor, generation, and a ledger of the records in the overlap):
Loads the checkpoint, or initialises it at now − window on the first run or when the environment changed (
checkpoint_reseton the receipt).Loops windows
[cursor − overlap, min(now − 10 s, cursor + window)]until it catches up or the run budget (activity_max_seconds) is nearly spent — at most 200 windows a run.Runs one bounded audit search per window (both the legacy
key=valueand the JSON audit formats,action=searchcandidates), pages the results, normalises every row through the evidence adapters and correlates the search ids.A window over
activity_max_recordsis halved down to 60 seconds; at the minimum the bounded evidence is kept and the cursor advances with an explicit gap (record_limit_exceeded_at_minimum_window) — a coverage gap, never a stall.The overlap re-reads the last
activity_overlap_secondsof ingestion time so a late event is never missed; a record the ledger already holds is skipped, so nothing is written twice.A blocking diagnostic (transport, search failure, incomplete pages) stops the run without advancing; the window is retried on the next runs, up to
activity_max_window_retries, after which its partial evidence is kept, the cursor advances andwindow_retries_exhaustedis recorded.Evidence-quality diagnostics (
normalization_gap,locator_conflict,invalid_event_envelope,unsupported_source_contract) are recorded on the receipt but never block: the observations are indexed with their status.The batch is acknowledged, its events published, and the checkpoint persisted after every window.
After the audit windows, the same run collects the resource usage (_introspection
PerProcess search samples, one row per search id and account with CPU seconds
integrated from the sampled percentages, the part spent on the search peers, the peak
memory, elapsed time, node count and Splunk’s own type / provenance / mode / label —
windows up to now − 120 s to absorb the introspection lag, a failed window retried, an
overflowing one keeping its most expensive searches with record_limit_exceeded),
refreshes the capacity at most hourly (the cores and memory of the search hosts: the
nodes that ran a meaningful share of the search processes, any role; a utility node
running its own few searches is listed but outside the basis), and scans the span it
advanced through for the logins and the automation evidence of the classification
(bounded, fail-open). The receipt carries the windows, the counts (observations written
/ folded / REST, executions, replayed), the complete flag, resources,
capacity, logins, automation and the detected accounts.
Source versions. The audit source contract is the line format, not the Splunk
version: a line whose sourcetype and action the reader knows is read whatever version
emitted it. The connected head is verified from its server info and its search peers
discovered once per run (when the account holds dispatch_rest_to_indexers); every
other emitter is parsed under the head’s version when source_version_assumption is on
(default), its record labelled assumed and the receipt carrying
source_version_assumed. A version family the reader does not know yet is read all the
same and flagged with the soft diagnostic source_version_unqualified, which never
blocks collection.
Policies¶
The policies job first reads the present saved-search rows of every tier under the
inventory lock (a bounded 90-second wait — if a long inventory run is still publishing, the
run fails with inventory_publication_in_progress and retries at its next schedule) and
evaluates the five inventory policies with the completeness of the latest inventory
receipts. When at least one activity or resource policy is enabled, it aggregates the
indexed trackme:uam:observation events over the trailing activity_policy_window_seconds
(one bounded stats per rule, one row per search first — a search’s grant and completion
never count twice) and one resource rollup (stats … by account, sid, the 20,000 most
expensive searches), and evaluates them per account. It then reconciles the findings
collection under the findings lock (see Findings lifecycle), expires
the temporary dismissals, publishes the finding totals the card reads, writes a
trackme:uam:finding event per transition, and finally drains the notification
outbox (fail-open, never a failed run).
The receipt stores one evaluation per policy (family, enabled, state, evaluated /
matched / excluded / unknown / not-applicable counts, readiness and thresholds source) plus
the activity window and its completeness; it is complete only when both families were.
Completeness is per family: inventory findings are unobserved only when every tier’s
inventory is complete, audit findings when the activity window was complete, resource
findings when the resource coverage is complete on top; a rule whose aggregation failed is
incomplete and neither creates nor closes anything. The bounded insight of each
evaluation (why records were unknown, excluded or not applicable, the owners and apps
behind each bucket, samples) is kept in the KV receipt only and served by
policy_evaluation.
Backfill¶
The backfill report collects the audit and resource evidence of the backfill_days
before the live activity start, through the route with the same failover as the activity
job. Evidence in an index cannot be taken back, so the job is built to never collect a
span twice:
The plan ends exactly at the activity checkpoint’s history start (the first live window’s start) and waits while it is unknown (
backfill_waits_for_first_activity_run). Resource buckets are clipped to the resource checkpoint’s own start.Fixed buckets (
backfill_bucket_seconds, 15 minutes by default) are walked backwards, newest history first, one bounded audit search and one resource search each; a bucket overbackfill_max_recordsis split in two.One attempt per run: a bucket with a blocking diagnostic is queued for the next hourly run and the run stops (a transient outage costs an hour, not the range); after
backfill_max_retriesattempts it is recorded inskipped. The run stops cleanly atbackfill_max_seconds(3,000 s by default, measured from the job start so the hour always covers the run). A bucket interrupted between its write and its commit is recorded as covered (interrupted), not read again.Retention edges come from the sources: one
tstats min(_time)per source at plan time; both floors complete the plan early (retention_reached). The introspection retention is typically shorter (about two weeks) than the audit retention, so the resource history may end earlier.Re-plans carry the coverage forward: a changed range or bucket size, or a Plan again, plans a new generation starting where the previous one stopped; a range already covered completes at once. A complete or cancelled plan un-schedules the report;
backfill_days=0disables it.The backfill also scans each bucket for the logins and the automation evidence, so the people and the services already working there when the tenant was created are classified without waiting.
The receipt carries the buckets done / split / retried / skipped / interrupted, the events
written, the progress (percent, ETA from the measured rate, per-source edges) and the
plan status; the tenant card reads it while the plan runs.
AI Findings Advisor¶
The AI Findings Advisor reviews the open and investigating findings of a UAM tenant —
or one finding — and, in act mode, records the decisions the evidence supports (status,
severity, assignee, a rationale on the history entry marked via AI Advisor; never the
investigation note). Interactive from the Findings tab toolbar (a sweep) and the Investigate
view (one finding), shown when AI is enabled; tenant administrators only. The scheduled
review is off by default: ai_uam_advisor_enabled in Manage AI Agents schedules
trackme_uam_advisor_tracker_tenant_<tid> (daily between 18:00 and 21:00), which runs one
advisor per finding not reviewed in the last ai_uam_advisor_min_days_between_reviews
days, within ai_uam_advisor_max_runtime_sec; the shared automated filters apply to the
finding severity. The report is owned by the tenant owner, who must administer the tenant.
See The AI Advisors.
Receipts and diagnostics¶
A receipt is value-free: counts, statuses, diagnostic codes and window bounds, never an
account name, a search text or a credential. The Collection health tab renders the
latest receipt of every job and, through the run view, the last seven days of runs from the
indexed receipts with each diagnostic explained. Common codes: checkpoint_reset,
record_limit_exceeded, window_retries_exhausted, source_version_assumed,
normalization_gap, inventory_publication_in_progress, capacity_unknown,
resource_collection_disabled, login_collection_failed, automation_collection_failed,
notification_drain_failed, retention_reached. A completed run with zero
observations is a successful run: absence of activity is not a failure. A completed run is
not a lossless-ingestion guarantee — the receipt says what was covered.
The technical logs are distinct from the evidence and carry outcomes only:
trackme_uam_tracker.log (sourcetype=trackme:custom_commands:trackmeuamtracker) and
trackme_rest_api_uam.log (sourcetype=trackme:rest_api), both with tenant_id=
index-time extracted, so index=_internal sourcetype=trackme:custom_commands:trackmeuam*
tenant_id=<tid> answers the usual support questions without touching the evidence.
The evidence¶
Events are appended to the dedicated trackme_uam_events.log under
$SPLUNK_HOME/var/log/splunk — its own file because UAM writes one line per search
(gigabytes a day on a busy deployment, against under a megabyte for every other component)
and must not share the rotating state-events ring of the other components — with the
target index, sourcetype and source (uam:<job>:<tenant_id>); Splunk’s default monitor
picks the file up and TrackMe’s ingest evaluation routes each event.
Sourcetype |
Family |
Content |
|---|---|---|
|
evidence |
One audit record per search and per window: |
|
evidence |
One correlation candidate per search id: the lifecycle (granted, completed…), the claims, the phases, the accounts and apps reported. |
|
evidence |
One search’s resource usage in one window: |
|
evidence |
One inventory change: the object’s fields, the |
|
findings |
One finding transition with the evidence and the investigation status. |
|
receipts |
The outcome of one tracker run. |
Every event carries tenant_id (an indexed field), environment_id,
event_kind and _time, plus indexed copies uam_kind, uam_environment,
uam_account, uam_sid, uam_app and uam_status so the per-account and
timeline reads run through tstats — O(accounts) instead of O(events) — with a raw
fallback for evidence indexed before an upgrade. The evidence and receipts families go to
the tenant’s summary index, the findings to its notable index; each family can be
pointed at another index (indexes.evidence / findings / receipts).
The compact record. Observations were the bulk of the metered volume (about 17 audit
lines per search on a lab tenant, two thirds of them REST access). The writer therefore
keeps one record per search per window: the completion record is written and the grant
folds into it (granted_time); the granted REST: lines (job lookups, status
polling, result streams) are not written at all — each account that accessed jobs over
REST gets one ``rest_access`` observation per batch with its call counts per kind
(job_poll, job_lookup, job_dispatch, stream, remote, other), so the
account keeps its presence and last activity; every null or unread key is dropped. With Store SPL search on, the search text adds
about a third to a record (≈ 500 bytes).
License impact. The evidence is indexed data and counts against the Splunk license
like every TrackMe summary event. The volume follows the search activity of the monitored
environment — one compact record per search and window, one resource record per search
and window, the inventory changes and one receipt per run — and is bounded by the
collection limits (Tenant options). Switch off
resources_enabled to drop the resource records (the resource policies then stay
learning), keep store_search_text off unless you need the SPL in the index (it can
carry literals and secrets, and it is readable by every TrackMe role of the tenant), and
size backfill_days with the retention of your sources in mind.
Identity and classification¶
The account key. Every observation carries account and account_source
(legacy_search_user, structured_actor or unknown). The legacy user header
wins when present, because in the field it names the caller while the structured actor is
often the system principal. It is a labelled source claim, not a verified initiator:
UAM never claims that a person sat at the keyboard, and a scheduled search runs under its
owner or the system account.
Classes and detection.
Class |
Detected when |
Limits |
|---|---|---|
|
An interactive search — a resource record whose provenance starts with |
Ad hoc alone is not a signal (connectors dispatch ad hoc searches through the API);
a labelled dispatch is not either (a dashboard set to run as its owner). Users
working only through React apps (TrackMe’s UI included) reach splunkd as
|
|
Sustained automation and nothing a person does: automation ( |
The automation days come from the |
|
Seeded for |
Assign it by hand to the tenant owner and to the administrator accounts that run the platform’s own workload. |
Rules. Up to 100 ordered classification_rules, each with a stable id, a name,
enabled, match (exact or wildcard — *, ?, [abc], [!abc] on
the whole name; no regular expression), pattern, case_sensitive and the
account_class. No naming heuristic runs unless an administrator saves a rule. Manual
account_classes (up to 10,000) always win.
Precedence, in every reader (the Users tab, the account view, the inventory, the policies): manual class → system default → first enabled matching rule → human detection → service detection. Removing a manual class returns the account to this automatic resolution; disabling a rule stops it at once; switching a detection off ignores its records without deleting them (a Splunk Web search run while the service detection was off still rules the account out when it is turned back on). Unmatched, undetected accounts stay unclassified. Settings → Account classification previews the draft (up to 1,000 names, the effective class and its source) before saving; the Users tab shows a detected marker with the evidence (signals, searches, logins, days) next to the class chip. Changing a class never rewrites historical evidence nor resolves an investigation: the policies use it on their next run.
The policy catalogue¶
Every policy has enabled, severity (low / medium / high /
critical), excluded_owners (inventory) or excluded accounts (activity),
excluded_apps (both default to Splunk_SA_CIM and trackme; up to 500 exact,
case-sensitive names, no wildcard) and its thresholds. Account policies add
account_scope: human (explicitly classified human — the default of the audit
rules), classified (human or service, never system — the default of the resource
rules) or all. Accounts outside the scope count as not applicable, excluded accounts
as excluded; excluded apps are removed from an account’s aggregate before the threshold is
compared. Every threshold is compared strictly (a finding needs more than the limit).
The server supplies the titles, logic, bounds and remediation (policy_catalog in the
workspace); the UI keeps no catalogue of its own.
Inventory policies (subject: a scheduled search)¶
The subject is one saved search of one tier; the finding id is stable across runs (policy version, environment, native record identity). Explicitly unscheduled or disabled searches are not applicable; a missing flag or an unreadable bound is unknown, never assumed. Evidence: the owner and its class (with the provenance of the class), the app, the tier, the schedule and the bounds read. UAM never edits nor dispatches the observed searches — the remediation is advisory.
Policy |
What it counts |
Thresholds, defaults and remediation |
|---|---|---|
|
Enabled schedules whose owner is explicitly human (unknown and service owners do not match). |
No threshold. On by default at low. Review whether this schedule should run under an approved service account, or record an owner exception. |
|
Enabled schedules whose owner has no human / service classification — a governance gap. |
No threshold. Off. Classify the owner in Settings after confirming the account’s purpose. |
|
Distinct configured minute slots in an active clock hour of a standard five-field cron (wildcards, steps, ranges, lists; duplicates count once; named weekdays, macros and extended forms are unknown). |
|
|
Enabled schedules whose |
No threshold. On. Configure a bounded dispatch time range where appropriate. |
|
Enabled schedules with an explicit real-time bound ( |
No threshold. On. Review whether a scheduled historical search meets the requirement. |
Activity policies (subject: an account)¶
Aggregated from the trackme:uam:observation events over the trailing
activity_policy_window_seconds (15 minutes), one row per search first, then per account
and app. The counts are distinct search ids the activity tracker observed — a
launch-count proxy, not independently verified launches — labelled with the coverage
(complete / partial). One finding per policy, environment and account: a repeated
burst updates the row. Evidence: the window, the count, the thresholds, the account class
and scope, the top five apps and, for costly_search, the maximum and total run time.
Policy |
What it counts |
Thresholds, defaults and remediation |
|---|---|---|
|
Distinct searches by one account in the window. |
|
|
Distinct searches whose reported |
|
|
Distinct completed searches whose reported earliest bound is explicitly unbounded
(legacy |
|
Resource policies (subject: an account)¶
Read from the trackme:uam:resource records (search heads and indexers, ad hoc and
scheduled) through one bounded rollup per policies run: CPU adds up across windows,
memory keeps its peak. CPU seconds are integrated from sampled percentages — CPU seconds on
any node, not SVC. The system class is outside the default classified scope (data-
model acceleration never becomes a finding by default) but its CPU stays in the share
denominator: it is real load.
Budgets from the capacity. With auto_thresholds on (the default), the absolute
limits are derived from the capacity record the activity job refreshes hourly —
search_cores (the core sum of the search hosts) and min_search_mem_mb (the
smallest of their memories) — and clamped to the threshold’s own range:
Policy |
Percentage (default) |
Derives |
|---|---|---|
|
|
|
|
|
|
|
|
|
With auto_thresholds off the stored absolute values apply. With it on but no usable
capacity, the stored values apply and the evaluation says manual_fallback. Evidence
carries thresholds_source (auto / manual / manual_fallback) and the
capacity basis; the policy modal shows the effective values. Readiness: the resource
rules are learning — nothing evaluated, nothing created, nothing closed — until the
tenant holds activity_min_history_hours (24) of resource history (the backfill counts)
and a capacity record.
Policy |
What it counts |
Thresholds, defaults and remediation |
|---|---|---|
|
The account’s CPU seconds over the window against its allowance. Its share of
the environment’s search CPU ( |
|
|
Distinct searches whose CPU seconds or peak memory reached the per-search minimum. |
|
|
Today’s accumulated CPU (UTC day) against |
|
Note
Why three resource angles. An earlier catalogue carried four overlapping resource
rules that opened two or three findings for one runaway account and asked operators
for a dozen absolute thresholds. The beta keeps three clear questions — volume per
account, intensity per search, behaviour against a baseline — with the share as a
severity boost rather than a finding, so another account’s load never opens and closes
findings. A configuration stored under the retired ids is converted on read and on
save, and their open findings are resolved by the policy (policy_retired) once the
replacement rule has evaluated fully.
Findings lifecycle¶
A finding is one row keyed by (policy, environment, subject) with status,
severity, policy_severity and severity_source (policy / manual),
assignee, note, observed, first_seen / last_seen / updated_at,
episodes, reopened_count, dismissed_until, a bounded history (the 50 most
recent records: {at, by, kind, from, to, note}, kind in status / severity / assignee
/ note for decisions, reobserved / unobserved / reopened / expired by the
policy) and the evidence.
Transition |
Meaning |
|---|---|
|
First match; inserted with |
|
Severity or evidence changed on reobservation (a threshold edit updates the evidence without changing the identity). |
|
Matched again after an absence: |
|
Not matched by a complete evaluation: the row is kept, |
|
A resolved finding observed again: |
|
A temporary dismissal ended: back to |
|
An investigation decision through REST or the UI; investigation fields are never overwritten by the tracker. |
|
The policy no longer exists and its replacement evaluated fully: resolved by the system, no notification. |
Decisions. A decision is any combination of a status (open, investigating,
resolved, dismissed), a severity override (low to critical, or
policy to restore the policy’s — possibly raised — value; the tracker keeps a manual
severity across runs), an assignee (a Splunk user; "" unassigns; the filters know
assigned to me and unassigned) and a note (omitted = unchanged, "" clears; it
rides on the status record when both change). Only a changed dimension leaves a history
record. Decisions apply to one finding, a selection (up to 500) or every open /
investigating finding of selected accounts (bounded to 20,000, refused above — never
truncated).
Resolve versus dismiss. Resolve says the investigation concluded: the policy
reopens the finding the next time it observes the condition (fifteen minutes later at
most). Dismiss says the finding is accepted: the policy never reopens it while it stays
dismissed, and its episodes keep counting. A dismissal can be temporary
(dismiss_for_seconds, one hour to one year): the finding returns to open once the
deadline passes; any status change clears the deadline, a new duration replaces it.
Dismissed findings sit behind the Dismissed tile (with the temporary count) and are
reopened from the row menu or the Investigate view.
Notes. Findings reuse TrackMe’s shared Notes (Markdown, permanent or expiring), keyed by the finding id, separate from the investigation note recorded in the history: context, not a decision.
Notifications¶
UAM has its own state-aware notifications, distinct from stateful alerting
(entity-centric) and Topology Alerts (view-centric): the finding is the alerting object,
and a notification is a consumer of the transitions the runtime already records — there is
no detector to configure and no “is this still alerting?” question, the finding row is
the state. One tenant-level setting (config.notifications, Settings → Notifications):
enabled(default off);eventsamongnew,reopened,escalated,resolved(on by default),dismissed,assigned(off);min_severity(medium);policies(empty = all);email_account(a TrackMe email delivery account, see Configuration; empty = the search head’s local mail transport);recipients(up to 50, filtered by the account’s allowed domains);mode:digestorper_finding.Producers: the policies tracker queues
new,reopenedandescalated(a policy severity raised — the share boost included — never over a manual override); a decision queuesresolved,dismissed,escalated(manual raise) andassigned.reobserved/unobserved/investigatingnever notify.Delivery: at the end of every policies run the outbox is drained (fail-open): rows older than 24 hours expire, rows the configuration filters out are suppressed, then one digest per run (grouped by event then severity, threaded per tenant and environment) or one email per transition threaded per finding (
Message-ID/In-Reply-To/Referencesderived from the finding id, so a mail client shows the finding’s evolution as one conversation). A failed send keeps the rows for the next run until they expire; the receipt carriesnotifications: {queued, sent, emails, suppressed, expired, failed, pending}.Every email carries the environment, the policy title, the subject block, the severity, the status and assignee, the last history entries, the evidence summary and a deep link that opens the Investigate view of the finding (
TenantHome?tenant_id=…&component=uam&tab=findings&finding=<id>).The indexed
trackme:uam:findingevents are the third channel: any Splunk alert, SOAR playbook or correlation search can consume them today.
Operations¶
Collection health. The tab shows, per tracker, the last run and its explanation, the
diagnostics with their meaning, the seven-day run history, the activity checkpoint (cursor,
lag, generation), the inventory datasets per tier, the audit windows of the last run, the
capacity, and the backfill plan. A scheduled run is invisible to the page (its lock never
disables the workspace; the 30-second timer notices its receipt landing); a Collect now
dispatched from the page is announced, disables writes while it runs and ends with a
Collection finished toast. A Collect now refused because a scheduled run holds the lock
(uam_job_running_or_recovery_required) is explained, not reported as an error.
Recovery. Locks carry the local server GUID and the search id. A lock at least five
minutes old whose search job has stopped on the same member is recovered automatically
at the next acquisition; Recover stopped run (POST recover) does the same on demand.
Locks of another member, legacy locks without a search id and configuration locks without
job ownership are never recovered automatically — verify the originating process on its
search head. The tenant lifecycle (delete, enable, disable, RBAC) has its own guard with
process ownership, so a request whose process died never blocks the tenant for good.
Guardian. The uam_collection_health check (tenant scope, warning, see
Configuration Guardian) runs at every tenant health-tracker cycle, reconciles the schedule
gate (tenant enabled, scheduled on, an owner set) and assesses the collection:
Problem |
Condition |
|---|---|
|
A tracker report is absent or un-scheduled while the tenant expects it. |
|
No receipt, or the last run finished later than its schedule allows (with a grace period). |
|
Three consecutive failed runs of one tracker ( |
|
The activity cursor did not advance past three schedules plus one window. |
|
A tier’s full listing is later than its cadence allows, or the last one did not complete within its budget. |
|
Notifications are enabled and the last drain failed to deliver. |
backfill stopped |
A backfill plan whose run stopped unexpectedly. |
The alert names the tenant, lists the problems in its metadata and points at the Collection health tab; it clears itself once the conditions are gone. Zero findings is a successful collection; unreadable inputs preserve an existing alert rather than clearing it.
Backfill controls. Pause stops after the current bucket and keeps the progress;
Resume continues at the next hourly run; Cancel retires the report (buckets already
collected stay indexed); Plan again plans the configured range as a new generation that
starts where the previous one stopped; Run now dispatches one slot immediately. Changing
backfill_days or backfill_bucket_seconds re-plans at the next run with the
coverage carried forward.
Lifecycle. Disabling the tenant un-schedules its reports; enabling reconciles them. Deleting a UAM tenant disables and deletes its reports, lookups and collections, cancels a running backfill, and refuses to run while a job lock is held (a stopped lock is recovered first). Repair UAM workspace (re-provisioning) re-posts the lookups with their field lists after an upgrade that added columns.
REST endpoints¶
Resource group uam (browse it with API & tooling → TrackMe REST API Reference; every
operation supports describe=true and requires tenant_id; tenant administrators
only). Every earliest / latest accepts what the Splunk time picker produces
(now, 0, -24h, -7d@d, an epoch); timeline, activity and accounts
echo the exact search they ran so Open in Search reproduces it.
Endpoint |
Purpose |
|---|---|
|
The workspace: configuration, policy catalogue, receipts per job key, locks, tiers, route and environments, checkpoint summary, counts, report presence, indexes. |
|
Paged findings with filters ( |
|
The Findings tiles and breakdowns over a range (one KV read, floors past 50,000 rows). |
|
One policy’s latest evaluation with its insight ( |
|
One policy’s counters per policies run over a range, from the indexed receipts. |
|
One tracker’s runs over a range ( |
|
Paged inventory rows per dataset across the tiers ( |
|
Bounded search of the evidence events of any kind, as the caller ( |
|
Evidence counts over time ( |
|
The Users view: accounts with class, roles, tiers, findings and activity counts
( |
|
One account’s profile, findings, timeline, recent searches, owned schedules and
resource usage ( |
|
The SPL of one search ( |
|
Validate, provision and reschedule with a full |
|
Evaluate a candidate |
|
Dispatch |
|
One investigation decision on one finding, up to 500 |
|
|
|
Remove a job lock older than five minutes whose search job has stopped
( |
|
Set or remove manual classes for 1 to 200 accounts ( |
The tenant is created with POST /trackme/v2/vtenants/admin/add_tenant carrying
tenant_uam_enabled=true and an optional uam_config object (any subset of the
options below — the wizard sends tiers, scheduled and backfill_days); see
Creating a splk-uam tenant.
Tenant options¶
The UAM configuration (uam_config at creation, config of configure afterwards,
edited from Settings). Every limit is a tunable with a validated range, not a product
cap.
Option |
Default |
Meaning |
|---|---|---|
|
|
The search head tiers (1–10): |
|
derived |
The route of the indexing-layer searches ( |
|
|
Schedule the trackers. |
|
|
Index the SPL of the reported searches with the audit records (Store SPL search). |
|
|
Parse emitters without a verified version under the head’s version, labelled
|
|
|
Collect the |
|
|
Detect human accounts from Splunk Web searches and browser logins. |
|
|
Detect service accounts from sustained automation. |
|
|
Audit window per chunk (60–3600). |
|
|
Ingestion-time overlap re-read per window (0–600, below the window). |
|
|
Audit records per window (100–200,000); overflow splits the window. |
|
|
Run budget of the activity job (30–900). |
|
|
Runs one window is retried on blocking diagnostics before it is skipped (1–20). |
|
|
Distinct searches per resource window (100–50,000); overflow keeps the most expensive. |
|
|
Resource history the resource policies need before they judge (0–720; 0 = none). |
|
|
REST page size of the inventory reads (50–500). |
|
|
Inventory records per run (100–1,000,000). |
|
|
Run budget of the inventory job (30–900). |
|
|
Hours between two full listings of a tier (1–168). |
|
|
Listing cadence while the tier’s write log cannot be read (1–24). |
|
|
Changed saved searches one delta run reads by exact path before listing instead (10–5000). |
|
|
Minutes between two reads of the users and roles (15–1440). |
|
|
Budget of one full listing (300–14,000). |
|
|
Start of the off-peak listing window, search head local time (0–23). |
|
|
Length of that window (1–24; 24 = any time). |
|
|
Findings the collection accepts (100–500,000); inserts stop with
|
|
|
Trailing window the activity and resource policies aggregate on (300–86,400). |
|
|
Days of history collected before the live start (0–90; 0 = no backfill). |
|
|
Span of one backfill bucket (300–3600). |
|
|
Wall-clock budget of one hourly backfill run (600–3300). |
|
|
Audit records per bucket (100–200,000); overflow splits the bucket. |
|
|
Attempts a bucket gets before it is recorded as a gap (1–20). |
|
|
Target index per family; empty = the tenant’s summary index (evidence, receipts) or notable index (findings). |
|
system defaults |
Manual classes per account ( |
|
|
The ordered assignment rules. |
|
the catalogue defaults |
Per policy: |
|
off |
See Notifications. |
The five ai_uam_advisor_* fields of the scheduled review live in the tenant’s
vtenant_account (Manage AI Agents), next to the shared automated filters.
Troubleshooting¶
Symptom |
What to check |
|---|---|
The inventory is complete but the Activity tab stays empty |
The route account cannot read |
Most audit lines are unparsed ( |
The emitters’ Splunk version could not be verified and
|
No service account is ever detected |
The |
The resource policies say learning for days |
Less than 24 hours of resource history ( |
A finding keeps coming back after Resolve |
Expected: the condition persists and the policy reopens a resolved finding. Fix the cause, record an owner / account exception on the policy, or dismiss it (temporarily or permanently). |
The Users tab shows a class the settings do not |
The class is detected (interactive activity or automation) — a manual class or a rule above it always wins; switch the detection off to ignore it. |
Collect now answers that a run is already in progress |
A scheduled run holds the lock; wait for its receipt. If no run is active, Recover stopped run in Collection health (after five minutes, same member only). |
The Guardian raises |
Open Collection health: the alert’s metadata lists the problems (a stale tracker, repeated failures, a stalled checkpoint, an overdue listing, a failed delivery). |
The backfill reports waiting for the first activity run |
Normal until the activity tracker has set the live start; it plans itself at its next hourly slot. |
The tenant cannot be deleted ( |
A tracker or backfill run is live or its lock is held by another member; wait for it or recover the lock on its member, then retry. |
Evidence of an old beta build is not tenant-filterable |
Events written before |
See also
UAM — User Activity Monitoring (Beta) — the overview.
Who is doing what on your Splunk? User Activity Monitoring (UAM) — the end-to-end walkthrough.
Creating a splk-uam tenant — the creation wizard and the REST flags.
Configuration Guardian — the Configuration Guardian.
Remote Splunk deployments — remote deployment accounts (the remote tiers).
The AI Advisors — the AI advisors and their scheduled reviews.