User Activity Monitoring — in depth

Tip

This is the in-depth reference for User Activity Monitoring (splk-uam, Beta in 2.4.18). For the readable overview of what UAM is and when to use it, start with UAM — User Activity Monitoring; for a guided, end-to-end walkthrough on a live deployment, read the User Activity Monitoring white paper. This page covers the tenant model, the search head tiers, the trackers and their receipts, the indexed evidence, identity and classification, the policy catalogue, the findings lifecycle, notifications, the operations, the REST endpoints, the tenant options and troubleshooting.

The tenant model

  • A UAM tenant is a dedicated tenant type: it carries tenant_uam_enabled and no entity component. Creation refuses any mix with an entity component (uam_tenant_must_be_dedicated), the Manage components menu is disabled on its card, replica trackers are refused, and the add / remove component endpoints answer uam_tenant_is_dedicated.

  • There are no entities: no decision maker, no impact scoring, no green / orange / red state, no stateful alerting, no priority / SLA / logical group. The unit of work is the finding; the alerting object is the finding’s transition.

  • Evidence lives in indexes, state in the KV store. Three per-tenant collections: kv_trackme_uam_settings_tenant_<tid> (the configuration, the checkpoints, the latest receipt and the lock of every job, the capacity, the detection records, the notification outbox), kv_trackme_uam_inventory_tenant_<tid> (one row per saved search / user / role, per tier) and kv_trackme_uam_findings_tenant_<tid> (one row per finding). Each has an administrative lookup trackme_uam_<kind>_tenant_<tid>.

  • Access is tenant-admin only: every UAM read and write requires trackmeadminoperations and membership in the tenant’s admin roles (or admin_all_objects). UAM evidence is audit data — the AI Assistant context of a UAM tenant is restricted the same way. Reads work on a read-only license; writes require a valid license beyond the Foundation edition.

  • The tenant card shows the open findings by severity (the Tenant Card Detail Level preference applies: compact mode shows open findings and the combined high / critical count), the tiers and the collection status; a double-click opens the read-only overview (collection strip, seven-day transitions timeline, tiles, findings table).

  • The tenant identifier of a UAM tenant is short — lowercase letters, digits and hyphens, up to 20 characters — because it keys the per-tenant collections and reports.

Search head tiers

One indexing layer, N search head tiers:

Data

Lives on

Collected

Saved searches, users, roles

each search head tier

inventory and listing, once per tier

_audit activity, _introspection resource usage, _internal

the indexing layer (every tier forwards there)

activity, resources, capacity, once per tenant, through the route

The Configure UAM step of the tenant wizard with two tiers — the local search head as the route account and a remote account tier that passed its connectivity check
  • tiers — 1 to 10 unique names, local and / or remote account names (see Remote Splunk deployments). Every selected account must feed the same indexing layer; that invariant is the administrator’s — one remote account per tier of that layer.

  • account — the route of the indexing-layer searches: local when selected, else the first tier; derived when a submission carries tiers without account.

  • Typical shapes: on Splunk Cloud, TrackMe on the ad-hoc search head (local) plus a remote account for the Enterprise Security search head; on Splunk Enterprise, any number of remote accounts over the same indexers; for central monitoring, remote accounts only.

  • A user present on several tiers is one account (roles united, tiers listed); a saved search belongs to its tier. Each tier has its own inventory report, lock and receipt, so a slow tier never starves another and a skipped tier is visible on its own.

  • A search head that keeps its internal indexes local (no forwarding of _audit / _introspection) hides its searches, logins and resource usage from the tenant — the forwarding prerequisite applies to every signal alike.

The trackers

Every job is a scheduled report running | trackmeuamtracker tenant_id=<tid> job=<job> as the tenant owner, with its own lock (never stolen; a lock whose search job is provably stopped is recovered after five minutes), its receipt (the outcome of the latest run, kept in the settings collection and indexed as trackme:uam:receipt) and its entry in the tenant’s operational status. Collect now in the workspace dispatches job=all (every tier’s inventory, then activity, then policies) as the caller.

Report

Job key

Schedule

trackme_uam_inventory_tracker_tenant_<tid>_tier_<account>

inventory:<tier>

*/15 * * * *

trackme_uam_listing_tracker_tenant_<tid>_tier_<account>

listing:<tier>

daily in the off-peak window (01:00 + 5 h, search head local time)

trackme_uam_activity_tracker_tenant_<tid>

activity

*/5 * * * *

trackme_uam_policies_tracker_tenant_<tid>

policies

3-59/15 * * * *

trackme_uam_backfill_tracker_tenant_<tid>

backfill

17 * * * * while a plan is open

trackme_uam_advisor_tracker_tenant_<tid>

(not a UAM job)

0 18-21 * * * when enabled

Inventory and listing

The inventory of a tier is the saved searches (with their owner, app, cron schedule, dispatch bounds, enablement and ACL namespace — SPL and alert actions are not read), the users and the roles, paged through TrackMe’s bounded REST reader (the local session or the remote account). A full listing costs the tier’s size whatever the filter, and an exact-path read is about ten times cheaper, so the collection is split:

  • inventory:<tier> (every 15 minutes) reads the changes reported by the tier’s own write log since the last run, by exact path (up to inventory_delta_max_objects — more escalate to a listing), and the users and roles every inventory_identity_minutes. While the write log cannot be read (_internal not searchable by the tier’s account, a member not forwarding it) the listing cadence falls back to inventory_fallback_reconcile_hours.

  • listing:<tier> runs the full listing every inventory_reconcile_hours inside the off-peak window, with its own time budget (inventory_listing_max_seconds); the first listing runs within the hour of the creation. List now dispatches one.

Each record is projected to a KV row under the tier’s environment and diffed against the previous row: a new row is added; changed field states changed (with changed_fields); a classification change alone reclassified; a row seen again after removal reappeared; a row absent from a complete scan is marked present=false with removed_at — never deleted. A partial or unavailable dataset upserts what was read and marks nothing removed. Change events are written only after the batch save succeeded, so a failed write never publishes a change twice. The receipt carries, per dataset, the status, source rows, pages, added / changed / removed / unchanged counts and diagnostics.

Activity

The activity job reads the tenant’s _audit trail from a durable ingestion-time checkpoint (cursor, generation, and a ledger of the records in the overlap):

  1. Loads the checkpoint, or initialises it at now − window on the first run or when the environment changed (checkpoint_reset on the receipt).

  2. Loops windows [cursor − overlap, min(now − 10 s, cursor + window)] until it catches up or the run budget (activity_max_seconds) is nearly spent — at most 200 windows a run.

  3. Runs one bounded audit search per window (both the legacy key=value and the JSON audit formats, action=search candidates), pages the results, normalises every row through the evidence adapters and correlates the search ids.

  4. A window over activity_max_records is halved down to 60 seconds; at the minimum the bounded evidence is kept and the cursor advances with an explicit gap (record_limit_exceeded_at_minimum_window) — a coverage gap, never a stall.

  5. The overlap re-reads the last activity_overlap_seconds of ingestion time so a late event is never missed; a record the ledger already holds is skipped, so nothing is written twice.

  6. A blocking diagnostic (transport, search failure, incomplete pages) stops the run without advancing; the window is retried on the next runs, up to activity_max_window_retries, after which its partial evidence is kept, the cursor advances and window_retries_exhausted is recorded.

  7. Evidence-quality diagnostics (normalization_gap, locator_conflict, invalid_event_envelope, unsupported_source_contract) are recorded on the receipt but never block: the observations are indexed with their status.

  8. The batch is acknowledged, its events published, and the checkpoint persisted after every window.

After the audit windows, the same run collects the resource usage (_introspection PerProcess search samples, one row per search id and account with CPU seconds integrated from the sampled percentages, the part spent on the search peers, the peak memory, elapsed time, node count and Splunk’s own type / provenance / mode / label — windows up to now − 120 s to absorb the introspection lag, a failed window retried, an overflowing one keeping its most expensive searches with record_limit_exceeded), refreshes the capacity at most hourly (the cores and memory of the search hosts: the nodes that ran a meaningful share of the search processes, any role; a utility node running its own few searches is listed but outside the basis), and scans the span it advanced through for the logins and the automation evidence of the classification (bounded, fail-open). The receipt carries the windows, the counts (observations written / folded / REST, executions, replayed), the complete flag, resources, capacity, logins, automation and the detected accounts.

Source versions. The audit source contract is the line format, not the Splunk version: a line whose sourcetype and action the reader knows is read whatever version emitted it. The connected head is verified from its server info and its search peers discovered once per run (when the account holds dispatch_rest_to_indexers); every other emitter is parsed under the head’s version when source_version_assumption is on (default), its record labelled assumed and the receipt carrying source_version_assumed. A version family the reader does not know yet is read all the same and flagged with the soft diagnostic source_version_unqualified, which never blocks collection.

Policies

The policies job first reads the present saved-search rows of every tier under the inventory lock (a bounded 90-second wait — if a long inventory run is still publishing, the run fails with inventory_publication_in_progress and retries at its next schedule) and evaluates the five inventory policies with the completeness of the latest inventory receipts. When at least one activity or resource policy is enabled, it aggregates the indexed trackme:uam:observation events over the trailing activity_policy_window_seconds (one bounded stats per rule, one row per search first — a search’s grant and completion never count twice) and one resource rollup (stats … by account, sid, the 20,000 most expensive searches), and evaluates them per account. It then reconciles the findings collection under the findings lock (see Findings lifecycle), expires the temporary dismissals, publishes the finding totals the card reads, writes a trackme:uam:finding event per transition, and finally drains the notification outbox (fail-open, never a failed run).

The receipt stores one evaluation per policy (family, enabled, state, evaluated / matched / excluded / unknown / not-applicable counts, readiness and thresholds source) plus the activity window and its completeness; it is complete only when both families were. Completeness is per family: inventory findings are unobserved only when every tier’s inventory is complete, audit findings when the activity window was complete, resource findings when the resource coverage is complete on top; a rule whose aggregation failed is incomplete and neither creates nor closes anything. The bounded insight of each evaluation (why records were unknown, excluded or not applicable, the owners and apps behind each bucket, samples) is kept in the KV receipt only and served by policy_evaluation.

Backfill

The backfill report collects the audit and resource evidence of the backfill_days before the live activity start, through the route with the same failover as the activity job. Evidence in an index cannot be taken back, so the job is built to never collect a span twice:

  • The plan ends exactly at the activity checkpoint’s history start (the first live window’s start) and waits while it is unknown (backfill_waits_for_first_activity_run). Resource buckets are clipped to the resource checkpoint’s own start.

  • Fixed buckets (backfill_bucket_seconds, 15 minutes by default) are walked backwards, newest history first, one bounded audit search and one resource search each; a bucket over backfill_max_records is split in two.

  • One attempt per run: a bucket with a blocking diagnostic is queued for the next hourly run and the run stops (a transient outage costs an hour, not the range); after backfill_max_retries attempts it is recorded in skipped. The run stops cleanly at backfill_max_seconds (3,000 s by default, measured from the job start so the hour always covers the run). A bucket interrupted between its write and its commit is recorded as covered (interrupted), not read again.

  • Retention edges come from the sources: one tstats min(_time) per source at plan time; both floors complete the plan early (retention_reached). The introspection retention is typically shorter (about two weeks) than the audit retention, so the resource history may end earlier.

  • Re-plans carry the coverage forward: a changed range or bucket size, or a Plan again, plans a new generation starting where the previous one stopped; a range already covered completes at once. A complete or cancelled plan un-schedules the report; backfill_days=0 disables it.

  • The backfill also scans each bucket for the logins and the automation evidence, so the people and the services already working there when the tenant was created are classified without waiting.

The receipt carries the buckets done / split / retried / skipped / interrupted, the events written, the progress (percent, ETA from the measured rate, per-source edges) and the plan status; the tenant card reads it while the plan runs.

AI Findings Advisor

The AI Findings Advisor reviews the open and investigating findings of a UAM tenant — or one finding — and, in act mode, records the decisions the evidence supports (status, severity, assignee, a rationale on the history entry marked via AI Advisor; never the investigation note). Interactive from the Findings tab toolbar (a sweep) and the Investigate view (one finding), shown when AI is enabled; tenant administrators only. The scheduled review is off by default: ai_uam_advisor_enabled in Manage AI Agents schedules trackme_uam_advisor_tracker_tenant_<tid> (daily between 18:00 and 21:00), which runs one advisor per finding not reviewed in the last ai_uam_advisor_min_days_between_reviews days, within ai_uam_advisor_max_runtime_sec; the shared automated filters apply to the finding severity. The report is owned by the tenant owner, who must administer the tenant. See The AI Advisors.

Receipts and diagnostics

A receipt is value-free: counts, statuses, diagnostic codes and window bounds, never an account name, a search text or a credential. The Collection health tab renders the latest receipt of every job and, through the run view, the last seven days of runs from the indexed receipts with each diagnostic explained. Common codes: checkpoint_reset, record_limit_exceeded, window_retries_exhausted, source_version_assumed, normalization_gap, inventory_publication_in_progress, capacity_unknown, resource_collection_disabled, login_collection_failed, automation_collection_failed, notification_drain_failed, retention_reached. A completed run with zero observations is a successful run: absence of activity is not a failure. A completed run is not a lossless-ingestion guarantee — the receipt says what was covered.

The technical logs are distinct from the evidence and carry outcomes only: trackme_uam_tracker.log (sourcetype=trackme:custom_commands:trackmeuamtracker) and trackme_rest_api_uam.log (sourcetype=trackme:rest_api), both with tenant_id= index-time extracted, so index=_internal sourcetype=trackme:custom_commands:trackmeuam* tenant_id=<tid> answers the usual support questions without touching the evidence.

The evidence

Events are appended to the dedicated trackme_uam_events.log under $SPLUNK_HOME/var/log/splunk — its own file because UAM writes one line per search (gigabytes a day on a busy deployment, against under a megabyte for every other component) and must not share the rotating state-events ring of the other components — with the target index, sourcetype and source (uam:<job>:<tenant_id>); Splunk’s default monitor picks the file up and TrackMe’s ingest evaluation routes each event.

Sourcetype

Family

Content

trackme:uam:observation

evidence

One audit record per search and per window: account and account_source, app, sid, search_type, the search status, the metrics (total_run_time, event_count, result_count, scan_count, exec_time), the time_bounds (search_et / search_lt as text, the N/A all-time marker included), granted_time, merged (what was folded into it) and the search text when Store SPL search is on. Also the rest_access summaries and the denied REST: lines (security evidence).

trackme:uam:execution

evidence

One correlation candidate per search id: the lifecycle (granted, completed…), the claims, the phases, the accounts and apps reported.

trackme:uam:resource

evidence

One search’s resource usage in one window: account, sid, app, search_type, provenance, mode, label, cpu_seconds, peer_cpu_seconds, peak_mem_mb, elapsed, nodes, samples, roles, window.

trackme:uam:inventory

evidence

One inventory change: the object’s fields, the change (added / changed / removed / reappeared / reclassified) and changed_fields.

trackme:uam:finding

findings

One finding transition with the evidence and the investigation status.

trackme:uam:receipt

receipts

The outcome of one tracker run.

Every event carries tenant_id (an indexed field), environment_id, event_kind and _time, plus indexed copies uam_kind, uam_environment, uam_account, uam_sid, uam_app and uam_status so the per-account and timeline reads run through tstats — O(accounts) instead of O(events) — with a raw fallback for evidence indexed before an upgrade. The evidence and receipts families go to the tenant’s summary index, the findings to its notable index; each family can be pointed at another index (indexes.evidence / findings / receipts).

The compact record. Observations were the bulk of the metered volume (about 17 audit lines per search on a lab tenant, two thirds of them REST access). The writer therefore keeps one record per search per window: the completion record is written and the grant folds into it (granted_time); the granted REST: lines (job lookups, status polling, result streams) are not written at all — each account that accessed jobs over REST gets one ``rest_access`` observation per batch with its call counts per kind (job_poll, job_lookup, job_dispatch, stream, remote, other), so the account keeps its presence and last activity; every null or unread key is dropped. With Store SPL search on, the search text adds about a third to a record (≈ 500 bytes).

License impact. The evidence is indexed data and counts against the Splunk license like every TrackMe summary event. The volume follows the search activity of the monitored environment — one compact record per search and window, one resource record per search and window, the inventory changes and one receipt per run — and is bounded by the collection limits (Tenant options). Switch off resources_enabled to drop the resource records (the resource policies then stay learning), keep store_search_text off unless you need the SPL in the index (it can carry literals and secrets, and it is readable by every TrackMe role of the tenant), and size backfill_days with the retention of your sources in mind.

Identity and classification

The account key. Every observation carries account and account_source (legacy_search_user, structured_actor or unknown). The legacy user header wins when present, because in the field it names the caller while the structured actor is often the system principal. It is a labelled source claim, not a verified initiator: UAM never claims that a person sat at the keyboard, and a scheduled search runs under its owner or the system account.

Classes and detection.

Class

Detected when

Limits

human

An interactive search — a resource record whose provenance starts with UI: (UI:Search, UI:Dashboard:<name>, UI:Report) and carries no saved-search label — or a successful browser login to Splunk Web (the audit login attempt events with a browser user agent). One is enough; the detection is sticky and never demotes. Switch: classification_detect_interactive.

Ad hoc alone is not a signal (connectors dispatch ad hoc searches through the API); a labelled dispatch is not either (a dashboard set to run as its owner). Users working only through React apps (TrackMe’s UI included) reach splunkd as rest:jobs — the login signal covers them.

service

Sustained automation and nothing a person does: automation (rest / rest:*, scheduler, summary_director, no provenance) on at least 7 distinct UTC days, never a person search nor an ambiguous one (a labelled Splunk Web dispatch, splunkjs, an MCP client, any unknown provenance), and not a detected human. Withdrawn as soon as an ambiguous search appears (the blocked count is sticky). Switch: classification_detect_service.

The automation days come from the _audit trail (provenance= / savedsearch_name= of the legacy key=value format): a JSON-only _audit gives no service detection, and the 7 days must lie inside the audit retention. An incomplete resource window counts no automation. Accounts with no activity are never detected.

system

Seeded for nobody, splunk-system-user and the built-in admin whenever unset; never detected, never demoted by a detection.

Assign it by hand to the tenant owner and to the administrator accounts that run the platform’s own workload.

Rules. Up to 100 ordered classification_rules, each with a stable id, a name, enabled, match (exact or wildcard — *, ?, [abc], [!abc] on the whole name; no regular expression), pattern, case_sensitive and the account_class. No naming heuristic runs unless an administrator saves a rule. Manual account_classes (up to 10,000) always win.

Precedence, in every reader (the Users tab, the account view, the inventory, the policies): manual class → system default → first enabled matching rule → human detection → service detection. Removing a manual class returns the account to this automatic resolution; disabling a rule stops it at once; switching a detection off ignores its records without deleting them (a Splunk Web search run while the service detection was off still rules the account out when it is turned back on). Unmatched, undetected accounts stay unclassified. Settings → Account classification previews the draft (up to 1,000 names, the effective class and its source) before saving; the Users tab shows a detected marker with the evidence (signals, searches, logins, days) next to the class chip. Changing a class never rewrites historical evidence nor resolves an investigation: the policies use it on their next run.

The policy catalogue

Every policy has enabled, severity (low / medium / high / critical), excluded_owners (inventory) or excluded accounts (activity), excluded_apps (both default to Splunk_SA_CIM and trackme; up to 500 exact, case-sensitive names, no wildcard) and its thresholds. Account policies add account_scope: human (explicitly classified human — the default of the audit rules), classified (human or service, never system — the default of the resource rules) or all. Accounts outside the scope count as not applicable, excluded accounts as excluded; excluded apps are removed from an account’s aggregate before the threshold is compared. Every threshold is compared strictly (a finding needs more than the limit). The server supplies the titles, logic, bounds and remediation (policy_catalog in the workspace); the UI keeps no catalogue of its own.

Activity policies (subject: an account)

Aggregated from the trackme:uam:observation events over the trailing activity_policy_window_seconds (15 minutes), one row per search first, then per account and app. The counts are distinct search ids the activity tracker observed — a launch-count proxy, not independently verified launches — labelled with the coverage (complete / partial). One finding per policy, environment and account: a repeated burst updates the row. Evidence: the window, the count, the thresholds, the account class and scope, the top five apps and, for costly_search, the maximum and total run time.

Policy

What it counts

Thresholds, defaults and remediation

launch_burst/v1

Distinct searches by one account in the window.

max_launches 1–100,000, default 200. Off. Confirm whether the burst is expected automation; raise the limit or record an account exception if it is.

costly_search/v1

Distinct searches whose reported total_run_time reached the minimum.

min_run_time_seconds 1–86,400 (300), max_costly_searches 0–100,000 (3). Off. Review the long searches (run time, scan count); work on time ranges, scheduling or acceleration.

all_time_search/v1

Distinct completed searches whose reported earliest bound is explicitly unbounded (legacy search_et=N/A, structured 0); a missing bound is unknown.

max_all_time_searches 0–100,000, default 0. On at high. Agree on bounded time ranges or an account exception for approved investigations.

Resource policies (subject: an account)

Read from the trackme:uam:resource records (search heads and indexers, ad hoc and scheduled) through one bounded rollup per policies run: CPU adds up across windows, memory keeps its peak. CPU seconds are integrated from sampled percentages — CPU seconds on any node, not SVC. The system class is outside the default classified scope (data- model acceleration never becomes a finding by default) but its CPU stays in the share denominator: it is real load.

Budgets from the capacity. With auto_thresholds on (the default), the absolute limits are derived from the capacity record the activity job refreshes hourly — search_cores (the core sum of the search hosts) and min_search_mem_mb (the smallest of their memories) — and clamped to the threshold’s own range:

Policy

Percentage (default)

Derives

resource_consumption/v1

cpu_capacity_pct (10 %)

max_cpu_seconds = search cores × window seconds × pct

resource_intensity/v1

search_cpu_capacity_pct (2 %)

min_cpu_seconds_per_search = search cores × window seconds × pct

resource_intensity/v1

search_memory_host_pct (10 %)

min_peak_mem_mb_per_search = smallest search host memory × pct

With auto_thresholds off the stored absolute values apply. With it on but no usable capacity, the stored values apply and the evaluation says manual_fallback. Evidence carries thresholds_source (auto / manual / manual_fallback) and the capacity basis; the policy modal shows the effective values. Readiness: the resource rules are learning — nothing evaluated, nothing created, nothing closed — until the tenant holds activity_min_history_hours (24) of resource history (the backfill counts) and a capacity record.

Policy

What it counts

Thresholds, defaults and remediation

resource_consumption/v1

The account’s CPU seconds over the window against its allowance. Its share of the environment’s search CPU (share_pct) is always in the evidence; strictly above dominant_share_pct the severity is raised one level and held for the rest of the observation episode. The share never opens a finding on its own.

max_cpu_seconds 1–10,000,000 (3600, derived), cpu_capacity_pct 1–100 (10), dominant_share_pct 1–100 (40). On, scope classified. Evidence adds the environment total, the heavy-search count and the top five searches and apps by cost. Open the account’s most expensive searches; work with the owner on time ranges, acceleration, scheduling or workload rules; raise the allowance or record an exception for approved workloads.

resource_intensity/v1

Distinct searches whose CPU seconds or peak memory reached the per-search minimum.

min_cpu_seconds_per_search 1–1,000,000 (300, derived), min_peak_mem_mb_per_search 64–1,000,000 (2048, derived), max_heavy_searches 0–100,000 (3), search_cpu_capacity_pct (2), search_memory_host_pct (10). On, scope classified. Evidence adds the heavy count and its CPU share. Review the heavy searches (type, provenance, label); agree on bounded time ranges, acceleration or scheduling.

resource_deviation/v1

Today’s accumulated CPU (UTC day) against max(multiplier × historical daily p95, floor), after at least seven fully covered prior days; missing coverage is unknown, never zero; today is never extrapolated.

lookback_days 7–30 (14), minimum_days 7–30 (7), multiplier 1–20 (3), the CPU floor (3600 s). Off, preview mode: matches are recorded, no finding and no notification. Review the daily baseline, the coverage and the expensive searches; confirm workload changes before adjusting the multiplier.

Note

Why three resource angles. An earlier catalogue carried four overlapping resource rules that opened two or three findings for one runaway account and asked operators for a dozen absolute thresholds. The beta keeps three clear questions — volume per account, intensity per search, behaviour against a baseline — with the share as a severity boost rather than a finding, so another account’s load never opens and closes findings. A configuration stored under the retired ids is converted on read and on save, and their open findings are resolved by the policy (policy_retired) once the replacement rule has evaluated fully.

Findings lifecycle

A finding is one row keyed by (policy, environment, subject) with status, severity, policy_severity and severity_source (policy / manual), assignee, note, observed, first_seen / last_seen / updated_at, episodes, reopened_count, dismissed_until, a bounded history (the 50 most recent records: {at, by, kind, from, to, note}, kind in status / severity / assignee / note for decisions, reobserved / unobserved / reopened / expired by the policy) and the evidence.

Transition

Meaning

new

First match; inserted with status=open.

updated

Severity or evidence changed on reobservation (a threshold edit updates the evidence without changing the identity).

reobserved

Matched again after an absence: episodes +1, a history record; open and investigating findings only gain the record.

unobserved

Not matched by a complete evaluation: the row is kept, observed=false. Absence is labelled, never auto-resolved; a partial evaluation never unobserves.

reopened

A resolved finding observed again: status back to open, reopened_count +1.

expired

A temporary dismissal ended: back to open at the next policies run, observed or not — the analyst asked for another look after that time.

triage

An investigation decision through REST or the UI; investigation fields are never overwritten by the tracker.

retired

The policy no longer exists and its replacement evaluated fully: resolved by the system, no notification.

Decisions. A decision is any combination of a status (open, investigating, resolved, dismissed), a severity override (low to critical, or policy to restore the policy’s — possibly raised — value; the tracker keeps a manual severity across runs), an assignee (a Splunk user; "" unassigns; the filters know assigned to me and unassigned) and a note (omitted = unchanged, "" clears; it rides on the status record when both change). Only a changed dimension leaves a history record. Decisions apply to one finding, a selection (up to 500) or every open / investigating finding of selected accounts (bounded to 20,000, refused above — never truncated).

Resolve versus dismiss. Resolve says the investigation concluded: the policy reopens the finding the next time it observes the condition (fifteen minutes later at most). Dismiss says the finding is accepted: the policy never reopens it while it stays dismissed, and its episodes keep counting. A dismissal can be temporary (dismiss_for_seconds, one hour to one year): the finding returns to open once the deadline passes; any status change clears the deadline, a new duration replaces it. Dismissed findings sit behind the Dismissed tile (with the temporary count) and are reopened from the row menu or the Investigate view.

Notes. Findings reuse TrackMe’s shared Notes (Markdown, permanent or expiring), keyed by the finding id, separate from the investigation note recorded in the history: context, not a decision.

Notifications

UAM has its own state-aware notifications, distinct from stateful alerting (entity-centric) and Topology Alerts (view-centric): the finding is the alerting object, and a notification is a consumer of the transitions the runtime already records — there is no detector to configure and no “is this still alerting?” question, the finding row is the state. One tenant-level setting (config.notifications, Settings → Notifications):

  • enabled (default off); events among new, reopened, escalated, resolved (on by default), dismissed, assigned (off); min_severity (medium); policies (empty = all); email_account (a TrackMe email delivery account, see Configuration; empty = the search head’s local mail transport); recipients (up to 50, filtered by the account’s allowed domains); mode: digest or per_finding.

  • Producers: the policies tracker queues new, reopened and escalated (a policy severity raised — the share boost included — never over a manual override); a decision queues resolved, dismissed, escalated (manual raise) and assigned. reobserved / unobserved / investigating never notify.

  • Delivery: at the end of every policies run the outbox is drained (fail-open): rows older than 24 hours expire, rows the configuration filters out are suppressed, then one digest per run (grouped by event then severity, threaded per tenant and environment) or one email per transition threaded per finding (Message-ID / In-Reply-To / References derived from the finding id, so a mail client shows the finding’s evolution as one conversation). A failed send keeps the rows for the next run until they expire; the receipt carries notifications: {queued, sent, emails, suppressed, expired, failed, pending}.

  • Every email carries the environment, the policy title, the subject block, the severity, the status and assignee, the last history entries, the evidence summary and a deep link that opens the Investigate view of the finding (TenantHome?tenant_id=…&component=uam&tab=findings&finding=<id>).

  • The indexed trackme:uam:finding events are the third channel: any Splunk alert, SOAR playbook or correlation search can consume them today.

Operations

Collection health. The tab shows, per tracker, the last run and its explanation, the diagnostics with their meaning, the seven-day run history, the activity checkpoint (cursor, lag, generation), the inventory datasets per tier, the audit windows of the last run, the capacity, and the backfill plan. A scheduled run is invisible to the page (its lock never disables the workspace; the 30-second timer notices its receipt landing); a Collect now dispatched from the page is announced, disables writes while it runs and ends with a Collection finished toast. A Collect now refused because a scheduled run holds the lock (uam_job_running_or_recovery_required) is explained, not reported as an error.

Recovery. Locks carry the local server GUID and the search id. A lock at least five minutes old whose search job has stopped on the same member is recovered automatically at the next acquisition; Recover stopped run (POST recover) does the same on demand. Locks of another member, legacy locks without a search id and configuration locks without job ownership are never recovered automatically — verify the originating process on its search head. The tenant lifecycle (delete, enable, disable, RBAC) has its own guard with process ownership, so a request whose process died never blocks the tenant for good.

Guardian. The uam_collection_health check (tenant scope, warning, see Configuration Guardian) runs at every tenant health-tracker cycle, reconciles the schedule gate (tenant enabled, scheduled on, an owner set) and assesses the collection:

Problem

Condition

report_missing / report_not_scheduled

A tracker report is absent or un-scheduled while the tenant expects it.

no_run_recorded / tracker_stale

No receipt, or the last run finished later than its schedule allows (with a grace period).

repeated_failed_runs

Three consecutive failed runs of one tracker (consecutive_failures, reset by a success or a changed environment).

checkpoint_stalled

The activity cursor did not advance past three schedules plus one window.

listing_overdue / listing_incomplete

A tier’s full listing is later than its cadence allows, or the last one did not complete within its budget.

notification_delivery_failed

Notifications are enabled and the last drain failed to deliver.

backfill stopped

A backfill plan whose run stopped unexpectedly.

The alert names the tenant, lists the problems in its metadata and points at the Collection health tab; it clears itself once the conditions are gone. Zero findings is a successful collection; unreadable inputs preserve an existing alert rather than clearing it.

Backfill controls. Pause stops after the current bucket and keeps the progress; Resume continues at the next hourly run; Cancel retires the report (buckets already collected stay indexed); Plan again plans the configured range as a new generation that starts where the previous one stopped; Run now dispatches one slot immediately. Changing backfill_days or backfill_bucket_seconds re-plans at the next run with the coverage carried forward.

Lifecycle. Disabling the tenant un-schedules its reports; enabling reconciles them. Deleting a UAM tenant disables and deletes its reports, lookups and collections, cancels a running backfill, and refuses to run while a job lock is held (a stopped lock is recovered first). Repair UAM workspace (re-provisioning) re-posts the lookups with their field lists after an upgrade that added columns.

REST endpoints

Resource group uam (browse it with API & tooling → TrackMe REST API Reference; every operation supports describe=true and requires tenant_id; tenant administrators only). Every earliest / latest accepts what the Splunk time picker produces (now, 0, -24h, -7d@d, an epoch); timeline, activity and accounts echo the exact search they ran so Open in Search reproduces it.

Endpoint

Purpose

GET /trackme/v2/uam/workspace

The workspace: configuration, policy catalogue, receipts per job key, locks, tiers, route and environments, checkpoint summary, counts, report presence, indexes.

POST /trackme/v2/uam/findings

Paged findings with filters (status, policy_id, severity, earliest / latest, owner, subject_type, assignee, observed, text, or one finding_id).

POST /trackme/v2/uam/finding_counts

The Findings tiles and breakdowns over a range (one KV read, floors past 50,000 rows).

POST /trackme/v2/uam/policy_evaluation

One policy’s latest evaluation with its insight (policy_id).

POST /trackme/v2/uam/policy_history

One policy’s counters per policies run over a range, from the indexed receipts.

POST /trackme/v2/uam/tracker_history

One tracker’s runs over a range (job_key): status, duration, coverage, events, diagnostics.

POST /trackme/v2/uam/inventory

Paged inventory rows per dataset across the tiers (dataset, scheduled_only, present, text).

POST /trackme/v2/uam/activity

Bounded search of the evidence events of any kind, as the caller (kind, account, app, sid, status, policy_id, transition, change, exclude_apps).

POST /trackme/v2/uam/timeline

Evidence counts over time (kind, span, split_by, account, app).

POST /trackme/v2/uam/accounts

The Users view: accounts with class, roles, tiers, findings and activity counts (text, classes, inventoried; truncated past 50,000 accounts).

POST /trackme/v2/uam/account

One account’s profile, findings, timeline, recent searches, owned schedules and resource usage (sections to select; concurrent bounded searches, incomplete names the ones that did not finish).

POST /trackme/v2/uam/search_text

The SPL of one search (sid): the indexed record when Store SPL search was on, else the environment’s own _audit record read on demand as the caller.

POST /trackme/v2/uam/configure

Validate, provision and reschedule with a full config; requires the current revision (optimistic concurrency).

POST /trackme/v2/uam/preview

Evaluate a candidate config against the current inventory — classes and policy findings — without writing anything (classification_only, sample_account).

POST /trackme/v2/uam/collect

Dispatch | trackmeuamtracker now (job: all / inventory / listing / activity / policies / backfill, optional tier).

POST /trackme/v2/uam/triage

One investigation decision on one finding, up to 500 finding_ids or the open findings of up to 500 accounts: status, severity, assignee, note, dismiss_for_seconds, via / rationale (advisor provenance).

POST /trackme/v2/uam/backfill

action: pause / resume / cancel / restart; answers the plan summary.

POST /trackme/v2/uam/recover

Remove a job lock older than five minutes whose search job has stopped (job: a job key).

POST /trackme/v2/uam/classify

Set or remove manual classes for 1 to 200 accounts (classes: human / service / system / "") without a full configure.

The tenant is created with POST /trackme/v2/vtenants/admin/add_tenant carrying tenant_uam_enabled=true and an optional uam_config object (any subset of the options below — the wizard sends tiers, scheduled and backfill_days); see Creating a splk-uam tenant.

Tenant options

The UAM configuration (uam_config at creation, config of configure afterwards, edited from Settings). Every limit is a tunable with a validated range, not a product cap.

Option

Default

Meaning

tiers

["local"]

The search head tiers (1–10): local and / or remote account names.

account

derived

The route of the indexing-layer searches (local when selected, else the first tier).

scheduled

true

Schedule the trackers.

store_search_text

false

Index the SPL of the reported searches with the audit records (Store SPL search).

source_version_assumption

true

Parse emitters without a verified version under the head’s version, labelled assumed.

resources_enabled

true

Collect the _introspection per-process search records.

classification_detect_interactive

true

Detect human accounts from Splunk Web searches and browser logins.

classification_detect_service

true

Detect service accounts from sustained automation.

activity_window_seconds

300

Audit window per chunk (60–3600).

activity_overlap_seconds

60

Ingestion-time overlap re-read per window (0–600, below the window).

activity_max_records

20000

Audit records per window (100–200,000); overflow splits the window.

activity_max_seconds

200

Run budget of the activity job (30–900).

activity_max_window_retries

3

Runs one window is retried on blocking diagnostics before it is skipped (1–20).

resources_max_records

5000

Distinct searches per resource window (100–50,000); overflow keeps the most expensive.

activity_min_history_hours

24

Resource history the resource policies need before they judge (0–720; 0 = none).

inventory_page_size

500

REST page size of the inventory reads (50–500).

inventory_max_records

100000

Inventory records per run (100–1,000,000).

inventory_max_seconds

200

Run budget of the inventory job (30–900).

inventory_reconcile_hours

24

Hours between two full listings of a tier (1–168).

inventory_fallback_reconcile_hours

4

Listing cadence while the tier’s write log cannot be read (1–24).

inventory_delta_max_objects

200

Changed saved searches one delta run reads by exact path before listing instead (10–5000).

inventory_identity_minutes

60

Minutes between two reads of the users and roles (15–1440).

inventory_listing_max_seconds

3600

Budget of one full listing (300–14,000).

inventory_listing_start_hour

1

Start of the off-peak listing window, search head local time (0–23).

inventory_listing_window_hours

5

Length of that window (1–24; 24 = any time).

findings_max_records

50000

Findings the collection accepts (100–500,000); inserts stop with findings_capacity_reached.

activity_policy_window_seconds

900

Trailing window the activity and resource policies aggregate on (300–86,400).

backfill_days

30

Days of history collected before the live start (0–90; 0 = no backfill).

backfill_bucket_seconds

900

Span of one backfill bucket (300–3600).

backfill_max_seconds

3000

Wall-clock budget of one hourly backfill run (600–3300).

backfill_max_records

20000

Audit records per bucket (100–200,000); overflow splits the bucket.

backfill_max_retries

3

Attempts a bucket gets before it is recorded as a gap (1–20).

indexes.evidence / indexes.findings / indexes.receipts

""

Target index per family; empty = the tenant’s summary index (evidence, receipts) or notable index (findings).

account_classes

system defaults

Manual classes per account (nobody, splunk-system-user and admin seeded as system).

classification_rules

[]

The ordered assignment rules.

policies

the catalogue defaults

Per policy: enabled, severity, exclusions, thresholds, account_scope, auto_thresholds.

notifications

off

See Notifications.

The five ai_uam_advisor_* fields of the scheduled review live in the tenant’s vtenant_account (Manage AI Agents), next to the shared automated filters.

Troubleshooting

Symptom

What to check

The inventory is complete but the Activity tab stays empty

The route account cannot read _audit on the indexing layer (roles on the remote deployment), or the search head keeps its internal indexes local. Check the activity receipt’s diagnostics in Collection health and run index=_audit action=search as the route account.

Most audit lines are unparsed (unsupported_source_contract)

The emitters’ Splunk version could not be verified and source_version_assumption is off; turn it on, or grant the route account dispatch_rest_to_indexers so the peers are discovered.

No service account is ever detected

The _audit trail is JSON-only (the automation scan reads the legacy key=value format), or its retention is below 7 days. Classify the accounts by rule or by hand.

The resource policies say learning for days

Less than 24 hours of resource history (resources_enabled off, or the introspection not forwarded), or no capacity record (the hourly refresh failed — see the activity receipt’s capacity). Lower activity_min_history_hours only if you accept budgets on a short history.

A finding keeps coming back after Resolve

Expected: the condition persists and the policy reopens a resolved finding. Fix the cause, record an owner / account exception on the policy, or dismiss it (temporarily or permanently).

The Users tab shows a class the settings do not

The class is detected (interactive activity or automation) — a manual class or a rule above it always wins; switch the detection off to ignore it.

Collect now answers that a run is already in progress

A scheduled run holds the lock; wait for its receipt. If no run is active, Recover stopped run in Collection health (after five minutes, same member only).

The Guardian raises uam_collection_health

Open Collection health: the alert’s metadata lists the problems (a stale tracker, repeated failures, a stalled checkpoint, an overdue listing, a failed delivery).

The backfill reports waiting for the first activity run

Normal until the activity tracker has set the live start; it plans itself at its next hourly slot.

The tenant cannot be deleted (uam_job_running_or_recovery_required)

A tracker or backfill run is live or its lock is held by another member; wait for it or recover the lock on its member, then retry.

Evidence of an old beta build is not tenant-filterable

Events written before tenant_id became an indexed field are left to age out; the workspace reads current evidence only.

See also