Release Notes
What is new and what was fixed in each release of CritterWatch, newest first.
If you are moving a running console between versions, read Upgrade Notes as well — it carries the changes that affect a deployed console's database schema, stored data, or monitored-host compatibility, and the manual steps for the ones CritterWatch cannot handle on its own. This page tells you what changed; that page tells you what you have to do about it.
1.1
Not yet released — 1.1.0-beta.8 is on NuGet now for testing; 1.1.0 follows once it has been exercised against real fleets. The first feature release after 1.0. It makes the console usable on fleets of hundreds of databases, turns the Event Model into a live screen rather than a diagram, lets a rebuild be scheduled for the middle of the night, puts recurring cron schedules on a screen you can pause from, lets the whole console run inside your own application, and opens the event store to AI agents. It also carries every fix from the unreleased 1.0.2.
Monitored services must move their store with the 1.1 client
The 1.1 client package depends on JasperFx 2.78, which added members every document store implements. A monitored service that takes the client must run Marten 9.43, Polecat 5.34 or Fisher 1.15 or later. With an older store the service restores and builds cleanly and then fails at startup with a TypeLoadException naming LoadDocumentAsync. See Upgrade Notes.
Upgrading from 1.1.0-beta.7
Telemetry stays under the broker's message-size limit on the largest fleets. A monitored service sends the console its telemetry as one message per push, and on a fleet of a few thousand tenants and hundreds of databases that message was approaching Amazon SQS's 256 KiB cap. Past it, the broker rejects the message and that service's projections and topology stop reaching the console with nothing on screen to say why. MaxItemsPerPush did not prevent this, because it counts items rather than bytes.
- Pushes are compressed harder: about 55–60% smaller on a large fleet's telemetry, for a few extra milliseconds per push on the monitored service.
- A push that would still be too large is sent as several. The new
MaxPushWireBytesoption (default 192 KiB) sets the limit. The console merges the pieces, so nothing is lost, and the service logs once, at Information, when it starts dividing. See Push size.
Both changes are in the monitored-service package and need no upgrade ordering: consoles and services on either side of them read each other's messages.
Also since beta.7:
- The console's 30-second refresh allocates less than half as much on a large fleet. Counting a service's agents by health no longer copies a string per agent: on a 75,000-agent service a quiet refresh now allocates about 9 MB instead of 20.
(Upgrading from beta.7 is a package bump on the console and on the monitored services, in either order.)
Upgrading from 1.1.0-beta.6
beta.7 is the large-fleet memory release. On a fleet of about 75,000 projection agents, one open browser tab could run the console out of memory: every connect built a 73 MB snapshot, half of it the agent roster, and wrote it as a single frame. The connect snapshot now carries five roster counts per service instead of the roster, and the Service Details and Node Detail pages load the agents they show when they open. Roster and agent-health updates go only to tabs that have one of those pages open.
The console also does far less work per telemetry cycle at that scale:
- A node's agent-health report costs about 150 reads for 75,000 agents instead of two round trips per agent. A healthy agent's row is no longer rewritten on every report, and its alert stream is read only when its health changed.
- The 30-second refresh reloads an agent roster only when it changed, and reads only the health rows and shard rows that moved. A quiet refresh tick at field scale allocates 46 MB rather than 273.
- Matching a projection shard to the node running it no longer scans the whole roster per shard.
Monitored services must also move their store with this client: see the warning at the top of this section.
Also since beta.6:
- The licensed service cap is enforced across the whole fleet. A console monitoring more services than its license allows now keeps the earliest registered (or the ones you choose with Keep monitored) and parks the rest: a parked service's telemetry is not applied and commands to it are refused, but its data is kept and it stays listed, tagged Not monitored — license limit. Before, only a brand-new service was checked, and a fleet already over the cap was monitored in full. See Licensing.
- Blue/green rollouts no longer read as lag. While two versions of a service ran side by side, each reported the other's projection shards as retired, the shard rows flipped between the two views, and the warming version's whole gap was reported as the service's worst lag. A retired report no longer overwrites a shard a running agent is reporting, and retired shards are left out of the health counts and the lag. The Operations shard panel says how many shards are no longer registered.
- The Document Explorer reads on the stores' shared contract: soft-deleted documents are hidden unless you ask for them (and then marked), All tenants reads every tenant explicitly, and reads are bounded (200 per page, a 256 KiB page budget, a 30-second timeout). Reading documents needs the new
documents.readcapability when permissions are wired, and every read is audited. See Upgrade Notes. - Per-tenant high-water marks carry their own liveness, so a partitioned store's tenants no longer inherit the staleness of a row the store has stopped writing.
- An idle service registers. A service with no telemetry to send, such as one with a document-only store, never completed its handshake with the console.
- The Event Explorer's Compact, stream fetch and query act in the selected tenant. On a store whose default tenant is disabled, a compact failed and an existing stream read as not found.
- Stream compaction compacts per tenant on a store whose default tenant is disabled.
- Recurring schedules have Trigger now, which runs a schedule once without changing its cadence.
- Listener health shows the in-flight limit and partition-lane depth, and the idempotency guard's dropped-duplicate count, when they matter.
- The Scheduled Messages page shows each message's tenant on a multi-tenant service.
- Per-tenant metrics alerts work on Prometheus and VictoriaMetrics sources.
- New MCP tool
list_tenants, and a refused hard delete of a disabled tenant now says why. - A mapping DDL fetch that fails says so, instead of reporting that no DDL was generated.
CRITTERWATCH_DISABLE_SQSlets a RabbitMQ-only development BFF start without LocalStack.- Built on WolverineFx 6.43.0, Marten 9.43.0, Polecat 5.34.0, Fisher 1.15.0, JasperFx 2.78.0 and Weasel 9.38.0. The monitored-service package,
Wolverine.CritterWatch, now depends on WolverineFx 6.43.0 and JasperFx.Events 2.78.0.
(No manual schema steps: the one new document, the license-slot pin, is created on first use. Upgrading from beta.6 is a package bump on the console, plus the store versions above on every monitored service.)
Upgrading from 1.1.0-beta.5
beta.6 fixes the Document Store Explorer on a service with more than one document store. It listed the document types of whichever store the service registered last, usually an ancillary one, for every store in the picker. A service's main-store types, whether registered up front or first used at runtime, never appeared, and browsing one of them returned an empty page instead of an error. Each store now answers for itself, and a request that names a store the service does not have, or names none on a service with several, gets an error listing the stores rather than a guess. The MCP document tools take an optional storeUri for the same reason. On Fisher, two document stores that share one SQLite file report the same identity and cannot be told apart (fisher#353): give each its own file.
Also since beta.5:
- Evicting a service no longer stops it from ever coming back. Eviction archived the service's streams, and every store refuses an append to an archived stream, so the service's next push failed on every restart. Its history is now deleted, so it registers again the next time it reports, and alerts it raised before can fire again. A service evicted by an earlier version needs one more eviction; see Upgrade Notes.
- A service whose every node has gone silent reads as Offline, with a new Critical Service offline alert. It only raises while the console is demonstrably receiving telemetry from other services, so it cannot fire on a console that monitors a single service. The Services page's health pill now reflects alerts raised by the console, not only ones reported by the services.
- Two nodes running the same projection shard now raise a Critical Shard ownership conflict, naming both nodes. It triggers no recovery, because restarting a shard two nodes are running cannot settle which one owns it.
- On a tenant-partitioned store, running per-tenant shards are no longer reported as orphaned. They were compared against a list that names projections without their tenant, so a healthy shard could read as retired and have its alerts suppressed.
- Per-tenant projection dead-letter counts reach the console. They read zero on every store.
- A failing database no longer empties the projection dead-letter page. The page shows every dead letter it could read and says how many databases could not be read.
- Stream compaction on a store whose default tenant is disabled now selects streams for a policy scoped to a tenant. A policy with no tenant scope on such a store says to declare
AllTenantsor aTenantId. Compacting per tenant, as opposed to selecting, waits on the next JasperFx release; an armed run on a named tenant reports that it cannot compact yet rather than failing per stream. - An empty new tenant database is no longer "needing attention" until it has events to process.
- The embedded console ships pre-generated handlers and never generates code at runtime. A Fisher console now creates its own event tables on a fresh database file; it could previously deadlock on its first write.
- EF Core conjoined tenancy is read from the
DbContext, and a dead letter from a conjoined tenant, replayed through the console, is handled for that tenant. - Built on WolverineFx 6.41.0, Marten 9.40.0, Polecat 5.32.0, Fisher 1.13.0, JasperFx 2.75.2 and Weasel 9.35.2. The monitored-service package,
Wolverine.CritterWatch, now depends on WolverineFx 6.41.0 and JasperFx.Events 2.75.2.
(No schema changes. Upgrading from beta.5 is a package bump, plus one step for any service an earlier version evicted.)
Upgrading from 1.1.0-beta.4
beta.5 fixes the embedded console, which did not work in a real application. Mounted with UseCritterWatchEmbedded(), it served its page at /critterwatch and then went nowhere. Every screen, the dashboard included, is loaded on demand, and the browser asked for those screens' files at your application's root (/assets/…) instead of under /critterwatch, so it got a 404 and navigation stopped. The header logo and the transport icons were broken the same way. Everything the console loads now resolves under wherever it is mounted, and it keeps routing when your application's Content-Security-Policy blocks inline scripts. This is now tested in a real browser against the packed package, on every release.
Also since beta.4:
- A browser tab that keeps losing its connection backs off properly. Reconnects that died right after reconnecting used to retry every second indefinitely, so one open tab on a very large fleet could keep the console rebuilding its connect snapshot. They now escalate to 30 seconds.
- The ingest-health chip's count follows the live number while a back-pressure alarm is raised, instead of repeating the count from the moment it was raised.
- On a tenant-partitioned store, shards are no longer reported as never started when their tenants have run them, and Inline and Live projections are no longer expected to have progress rows.
- Roll-up alerts stop writing a membership event every evaluation cycle when only the "N of M" denominator moved and the set of affected tenants did not.
- The Schedule Explorer shows a compaction policy as Failing as soon as the page loads, from the runs already recorded, rather than green until its next occurrence. Run now reports its outcome: the failure and its reason, or what a dry run would have compacted.
- The Audit Log shows a SQL statement as one line with its outcome, and the SQL view states its 1,000-row ceiling as the contract.
- A single-database store no longer leaves behind a frozen, database-less high-water row beside its live one. (A console upgraded from an earlier beta keeps the old row until its data is reset. It is harmless; see Upgrade Notes.)
(No schema changes, so upgrading from beta.4 is a package bump.)
Upgrading from 1.1.0-beta.3
beta.4 is the embedded-mode beta. The console now boots inside a Marten host on PostgreSQL queues with no broker, and that exact combination is under test for the first time — which found two things the docs now say plainly: the PostgreSQL transport needs role: MessageStoreRole.Ancillary or the host will not start, and the Marten registration takes an NpgsqlDataSource rather than a connection string (an earlier snippet had that wrong). Also since beta.3, four behaviours that had shipped without notes: the Conflicted projection verdict and its new wire field, an unknown high-water mark no longer reading as zero, the SQL allow-list resolving schemas (a bare name that would fall through search_path is now refused), and portable MCP tool schemas. The last two are behavioural — see Upgrade Notes.
Upgrading from 1.1.0-beta.2
beta.3 fixes three things a large multi-store fleet runs into immediately, all of them cases where a screen stated something it could not support:
- On a service with several event stores, the Event Store Explorer opened on whichever store the host happened to register first while the picker read "All stores". On one fleet that was an ancillary authorization store, so the Configuration tab described a 514-database service as having three event types and one projection. It now opens on the store the service calls its main one, and the picker names that store instead of claiming a scope the tab does not have.
- The Streams tab could load forever on a store with hundreds of databases — no timeout, no error, and the Refresh button disabled for the whole wait, so the one control that could retry was the one control you could not reach. Every Event Store Explorer read now has a 45-second deadline that turns a hang into an error you can retry, and Refresh stays clickable while a read is in flight.
- The rollout panel could fill the top of the screen with an answer it did not have — "Rollout in progress" over a dozen identical rows reading "Cannot tell", on a store where no shard can be measured at all. It now collapses to a single line saying so, and carries the age of the snapshot it is describing, because "in progress" is a present-tense claim.
(No schema or data changes, so upgrading from beta.2 is a package bump.)
Upgrading from 1.1.0-beta.1
beta.2 is worth taking if you run a large fleet, and especially if your console has ever looked calm while something was wrong. Most of Every screen says how old it is below landed after beta.1, including the one that matters most: a console that has stopped applying telemetry now says so instead of reporting healthy.
(beta.1 was published as 1.1.0-beta.1. An earlier draft of this page spelled it 1.1.0-beta-1, with a hyphen, which is not a version that exists on NuGet.)
Projections at fleet scale
The Operations screen was designed against a test fleet of two databases. A production console running 512 databases and 11,637 shards met it very differently, and much of this release is that gap.
- The shard is the row now, and the screen opens on the exceptions instead of on a complete inventory you have to search.
- Shard progression is queried server-side — paged and filtered, with lag derived where the rows already live. Connect used to ship every shard state to every browser tab: 13.7 MB, rebuilt client-side into three indexes costing 148.9 MB of heap. The live update was a broadcast too, so a dashboard tab held ~37 MB of shard state it never rendered.
- The Operations tab no longer blocks the browser. It spent 42–63 seconds in a single task on a 512-database store, because every derived lag reading re-walked all 11,637 shard states eight to ten times per row.
- Four projection-health signals were sampling one shard out of 512. A paused shard in a sharded fleet rendered "Healthy · Up to date".
- Alerts no longer go quiet exactly where they matter. A shard whose lag could not be computed was filtered out of the roll-up and raised nothing at all, and a shard that was never assigned an agent raised nothing because the check returned early on a null heartbeat. Relatedly, a projection whose lag cannot be measured no longer renders as "Healthy · Up to date" — twenty of twenty read that way while the Max Lag tile simultaneously called their gap unmeasurable.
- Per-tenant metric roll-ups are no longer broadcast to every tab every 30 seconds (2,173 rows per tick on one field fleet); that answer is requested on demand.
- Agent state is no longer republished when nothing changed — 2.76 MB of byte-identical state every 30 seconds, plus 145 heartbeat messages per tick.
- Dead-letter summarize stopped fanning out one message per message database — 518 messages every 30 seconds on every route, which drowned the 500-entry Message Log at 99.4% occupancy.
- A shard that two processes both believe they own is
Conflicted, notPaused. AProgressionProgressOutOfOrderExceptionstops a shard, and the store will never restart it on its own. It used to render as Paused — the one verdict that sends an operator to Restart, which cannot work while the second owner is still writing. The new verdict ranks above Failing and is answered by finding the other owner. It rides a new wire field; older satellites fall back to the exception type and reach the same verdict, so nothing degrades in the meantime. - An unknown high-water mark says so. A store with per-tenant partitioning has no store-global mark, and the Operations header used to render that as "High-Water Mark 0 — no events appended yet" directly above a grid showing shards at 4,403 / 1,109 / 70. Unknown is now reported as unknown, with the reason.
The Event Model is a live screen
The Event Model is now assembled from what your fleet actually does, rather than from a diagram someone drew.
- One semantic model, shared with the rest of the Critter Stack rather than rebuilt inside the console, merged from four sources — including what Wolverine observes at runtime.
- It renders fleet-wide with no traced subject. Previously the Workflow screen's Event Model mode showed nothing until you traced something, even though the merged model needs no subject.
- It refreshes when the model changes instead of being fetched once on mount and then quietly going stale as services connect.
- Where sources disagree, the screen says so rather than silently preferring one.
- A service that hosts several Event Models keeps all of them. Each of the three stores can name its own model, so a modular monolith with a store per module legitimately assembles several — and the console used to pick one name and fold the rest into it, losing a model's name outright with nothing reporting it. A service now reports the set, and where a surface can still carry only one model the fold is explicit and says so on the canvas, naming the models that went into it.
- An observed slice names the service it was watched on —
critterwatch://observed/{service}, and every service when two contribute to one slice — so a fleet-level disagreement names its sources and not only its rungs. - A projection's apply set reaches the canvas from the code, not only from a running store. The source generator now reads a projection class's
Apply/When/Create/ShouldDeletesignatures and contributes a View slice named after the document, carrying the events it consumes. It is named the way the store's own registry-derived slice is, so the two merge into one slice — and where the static reading and the running store disagree about what a projection applies, that is a hotspot on the canvas naming both, which is a finding rather than a conflict. - A Bobcat on a newer JasperFx than the console no longer empties the evidence. One vocabulary member the console had never heard of used to discard the whole spec-evidence fetch — the run and every scenario already retrieved. Unknown members now read as unrecognised, and a model that cannot be read keeps the run beside a note saying so.
- Drift is shown from evidence, not from a declaration — specification runs are read over HTTP, so a slice's colour reflects whether its behaviour is actually covered and passing.
Schedule a rebuild for tonight
- "Rebuild later…" on a projection row, defaulting to 03:00 tomorrow — the point of the feature is to move a rebuild out of business hours, so the default already is.
- The schedule lives in the monitored service's own message store, which means it appears on the Scheduled Messages page you already have, and can be cancelled or moved there. A console-held schedule would have been invisible and uncancellable.
- Acceptance is acknowledged immediately, carrying the time it will run, and the real outcome is acknowledged when it runs. A row no longer sits "pending" for hours against a rebuild that is correctly waiting for 03:00.
- Rebuild every tenant of a projection from one place. On a per-tenant partitioned store each tenant owns its own shard, so a store-global rebuild would have left every tenant partition untouched — the console now fans out per tenant and names the count before it does, because at fleet scale that is hundreds of simultaneous rebuilds.
- Scheduling, cancelling and rescheduling are available to AI agents through licensed MCP tools.
- A subscription rewind can be scheduled the same way. The Rewind dialog gains Now / Later, with every mode — to the beginning, a sequence or a timestamp. It lives in the same message store, shows on the same Scheduled Messages page, and a malformed target is refused when you schedule it rather than at 03:00. Requires the monitored service to be on this release of
Wolverine.CritterWatch; an older one has no handler for it and says so, rather than rewinding immediately.
WARNING
A schedule is only as durable as the queue it sits on. A service whose local queue is the default in-memory one loses a 03:00 rebuild to a 23:00 redeploy, silently. CritterWatch refuses a schedule it cannot honour rather than accepting one it will drop.
Schedules that repeat
Wolverine 6.34 can run a message on a cron schedule, and CritterWatch is the console for it.
- Register a schedule in one line —
opts.Schedules.ScheduleRecurring<NightlyRollup>("0 2 * * *"). Registration is the opt-in; there is no separate feature flag. - A fleet-wide Schedule Explorer lists every recurring schedule across every monitored service — name, cron expression, time zone, the message type each occurrence publishes, and when it next fires. Scoped to one service with the header selector, like the other explorers.
- Pause and resume a schedule from the console. Pausing stops it producing occurrences and cancels the one already pre-scheduled, so nothing fires in the gap. Both are separate permissions (
recurring-schedule.pause/.resume), because stopping production work and restoring it are not equally consequential. - The occurrences themselves were always visible. A pre-scheduled firing is an ordinary durable inbox row, so it appears on the Scheduled Messages page you already have. What was missing was the schedule — what fires, on what cadence, and when next.
- Available to AI agents through
list_recurring_schedules,pause_recurring_scheduleandresume_recurring_schedule.
WARNING
A pause is only durable if the service can store it. A service whose message store carries the recurring tracking table records the pause as a row that survives a restart. One without it applies the pause to the running agent's memory instead — real, but lost on the next restart. The console reports the difference rather than showing both as "paused", because a pause that a redeploy will quietly undo is worse than one that visibly failed.
TIP
Resuming does not backfill. Occurrences that would have fired while a schedule was paused are gone by design — the next occurrence is computed strictly after the resume. A missed firing stays a missed firing, which is the whole reason the schedule is owned by an agent rather than by a message that reschedules itself: one lost link cannot end the chain.
TIP
Schedules live on a service's main database and are never per-tenant, even on a multi-tenant service. The next-fire time and paused flag shown in the console are as of that service's last capability announcement; a pause or resume you perform from the console updates the row immediately from the acknowledgement.
Compact old streams on a schedule
Event streams grow forever. Compaction folds a stream's history into a single snapshot event — destructive and irreversible, but invisible from the read side, because the stream still aggregates to the same state and keeps its version. CritterWatch could already do that to one stream, by hand, from a button on the Event Store Explorer. That is not an answer for a fleet.
A compaction policy says it declaratively: compact any stream of this aggregate type whose un-compacted growth exceeds this many events, on this cron schedule.
monitoring.CompactStreams(p =>
{
p.Name = "BoardHistory";
p.AggregateType = typeof(Board);
p.GrowthThreshold = 1_000;
p.Cron = "0 2 * * *";
});- A policy is a recurring schedule, so everything the section above describes applies to it unchanged: one tick per cluster rather than one per node, a Schedule Explorer row, and pause and resume from the console under the same permissions.
- Growth is measured since the last compaction, not from zero. A threshold on the raw version would re-fire forever on a stream that has already been compacted; a threshold on growth reads zero on one that is fully compacted, which is what an operator means by "this stream is fine".
- Every run is reported on the Timeline — streams matched, compacted, skipped and failed, and why a run could not happen.
streamsMatchedis the whole backlog rather than the tick's share, so a policy working through a large store over successive nights is visible as progress. - Run one now from the Schedule Explorer, or over MCP, instead of waiting for the cron time.
- Available to AI agents through
list_compaction_policies,get_compaction_activity,run_compaction_policy, and pause and resume — the last three permission-gated, and the read tools explicit that a dry run changed nothing.
WARNING
A new policy is a dry run until you arm it, and arming it is a code change. A dry-run policy reports exactly what it would compact and deletes nothing. Arming is deliberately not an operator action: the console can stop a policy, run one, and show what one did, but it cannot turn a reporting policy into a destructive one — that stays in the code you deploy and review. The intended loop is declare, deploy, run it once from the console, read the report, then decide.
TIP
All three event stores compact — Marten, Polecat and Fisher — so a policy behaves the same whichever one a service runs on. Polecat needed 5.25.0 to get there: it implemented the typed compaction path but not the untyped one a policy needs, so it selected the right streams and then refused every one of them (polecat#572). Against an older Polecat a policy still reports that, once per run as a store-capability error rather than as a failure per stream.
TIP
Declaring your first policy creates a table. A policy is driven by a recurring schedule, and registering the first schedule is how a service opts in to the recurring-message agent — which provisions its tracking table on the service's main database. A service that declares no policy gets no schedule, no agent and no schema change, so adding the package alone never touches your database.
TIP
MaxStreamsPerTick caps how much one run does, and it exists for the first armed run against a store with a long backlog: the policy works through it over successive ticks rather than in one stampede. It matters most on SQLite, which takes a single writer per file.
Per-tenant subscription lag
Subscriptions are reported per tenant, not just per store. On a tenant-partitioned event store a subscription shard carries its own progression per tenant, and the console now reads and renders it — so a tenant that has stopped advancing is visible instead of averaged away.
TIP
A tenant with no progression row is a different thing from a tenant that is fully behind, and the console now distinguishes them. A missing row usually means the tenant has no registered partition, or a shard that started and never wrote — both are diagnosable, and neither is "caught up".
Sharper questions of the event store
The Event Store Explorer's by-metadata search and the query_events tool both gain the filters the underlying store query grew this cycle, and they AND together with the facets already there:
- Several event types at once, as a union rather than one query per type.
- A timestamp window and a sequence window, each inclusive and each valid half-open.
- Counting over a window — how many of this event between Tuesday and Thursday is now one question rather than a page-by-page tally.
- DCB tags combine with all of it. The Explorer's separate By Tag mode is gone: tags are one facet of the By Filters view, AND-ed with every other filter, with a total count of the combined match.
query_eventsdoes the same. - Tag queries work on every store. The old By Tag mode used a tag query that Fisher (SQLite) does not implement and Polecat (SQL Server) refuses store-wide, so on two of the three stores it answered with an error. The combined query is honoured by all three.
- By-Metadata results show their payloads. Every row used to render an empty body — the rows arrived in a different shape from the one the timeline reads. Fixed in the console, so it holds even against monitored services that have not upgraded.
- The Explorer opt-in now covers every event read.
EnableEventStoreExplorerused to gate the stream and tag views but not the by-metadata query orquery_events, which return full event bodies. It gates all of them now — a behaviour change for services that have not opted in; see the upgrade notes. - Listings answer once per database on a multi-database store. On a database-per-tenant or sharded store the Streams tab used to show whichever database the store's default session resolved, returned as though it were everything — 512 databases, one database's rows, nothing saying so. Recent streams and projection statuses now fan out over the store's databases and every row says which one it came from; a row opened from the listing reads its stream from that database, since a stream id is unique within a database and not across them. A single-database store keeps its one query. Needs a current client package on the monitored service. All three stores answer per database: SQL Server and SQLite always did, and PostgreSQL does from Marten 9.37 (marten#5383). The version that matters is the monitored service's own, not the console's — a service still on Marten 9.36 or earlier refuses the per-database read by name, per database, which remains an honest empty rather than one database's rows dressed as the whole store.
- A projection's daemon state is what the store actually knows. The three stores had each decided independently what a shard's state field meant; upstream now rules it a fact about the running daemon, with
Unknownmeaning "no daemon to ask" — which the console keeps distinct fromStopped, the reading an operator acts on and the one SQL Server used to report for it. Lag on a stopped daemon is real again too: the store head is the store head, not where the daemon got to, so a stopped projection no longer reads as zero behind at the moment it matters.
An offset-less timestamp is read as UTC, not as the console host's local zone. Applying a server's zone would shift a window by hours without saying so, and make two consoles disagree about the same store.
WARNING
The results you are shown are the results you asked for, or you are told they are not. A monitored service running an older client package does not understand a filter it has never heard of — it drops it while deserializing the request, so the filter never reaches the store to be refused, and the store's own guard against silently-ignored filters cannot see it. Unfiltered results would come back looking exactly like filtered ones.
So a service now echoes back the filters it actually applied, and both the console and the tool compare that against what was sent. If anything was dropped, the filters are named and no rows are rendered at all — an unfiltered page is worse than no page, because it reads as an answer. A service that echoes nothing is treated as having dropped everything, since the absence of an answer is not evidence that the answer was yes.
TIP
Tag queries combine with the filters above. A DCB tag match ANDs with event types, time windows and the rest, so "events tagged customer: acme, of this type, between Tuesday and Thursday" is a single call rather than two queries intersected by hand.
This landed late. The composable query object had no name/value tag member at all — the typed form resolves tag types against the store's own registered graph, which a console cannot reach, so a dictionary tag query could only go through a separate method that accepts no other filter. That gap is jasperfx#801, and closing it took all three event stores.
⚠️ A tags-only query against a monitored service too old to honour tag filters on the combined path has no totalCount. It is detected — its applied-filter echo lacks TagValues — and re-asked on the older standalone path, which that service does speak. Every other query, tags-only included, answers with a real total. Do not read a missing totalCount as zero.
AI and MCP
- Query the event store in SQL — the Event Explorer's new SQL view and the
query_sqlread tool run one read-onlySELECTon the monitored service, on its own connection, so the console never holds database credentials and a Marten, Polecat or Fisher service is reached the same way. The statement is allow-listed to the tables the store declares (a reference outside them is refused with the list), runs inside a read-only transaction that is always rolled back, is statement-timeout bounded and row-capped, and says when the cap truncated the answer rather than letting a floor read as a total. It is the escape hatch beside the structured query for the join, theGROUP BYand the per-stream count the filters cannot express. It carries its own capability (sql.queryin the console,mcp.sql.queryover MCP) and every run is written to the audit log — the outcome as well as the request, so a refused statement no longer reads as "Ran SQL". Opt-in on the monitored service:EnableSqlQuery, which followsEnableEventStoreExplorerunless set. ⚠️ The allow-list resolves schemas. An earlier build matched bare table names and ignored the schema, so one service's SQL view could read another service's event store in the same database. A bare reference is now qualified against the store's own schema and matched qualified-only; a reference that would have resolved elsewhere throughsearch_pathis refused. - Step a projection through its events as a read tool, alongside a diagnostic that explains what a projection did with a stream.
- Manage stream compaction policies — list what is declared, read what each run did, and run, pause or resume one. The read tools state plainly whether a run was a dry run that changed nothing, so an agent cannot report compaction that did not happen.
- Every tool remains licensed and capability-gated. The stepper and diagnostic only read. The compaction run, pause and resume tools act on the service: running a policy needs
compaction-policy.run, and pausing or resuming one needs the same capabilities as pausing or resuming a recurring schedule. - Every tool's input schema is portable. All 66 tools advertised schemas using
typeunions with"null", theformatkeyword and noadditionalProperties— the constructs MCP Inspector, OpenAI strict function-calling and Gemini's schema subset flag or reject. They now emit the portable subset, which is the difference between the MCP surface working with a given client and not.
Alert triage by a language model
Alerts that fire together are already coalesced into one batch. That batch can now be handed to a model, which writes a short assessment of what is probably going on and suggests what to do about it. The narrative is stored on each alert it names and rendered beside it in the console.
Off unless you turn it on. It calls a paid model on a schedule you do not control, and an alert storm is when it fires hardest — so it is a deliberate opt-in, the model client is yours to register (any Microsoft.Extensions.AI provider; CritterWatch takes no vendor dependency), and the token budget is enforced by a middleware that dead-letters an over-budget request rather than paying for it.
- One request per batch, never one per alert. The correlation is the feature: the model is shown the alerts that fired together as one possible incident, which is what an operator otherwise has to do in their head across five notifications.
- Suggestions mark the buttons the console already offers; they never add new ones. A suggested action lights up the existing remediation button for that alert type, and pressing it goes through the ordinary command path with the ordinary permission check.
- Nothing it returns is executed, and nothing it says grants anything. The narrative is labelled as model output on its face, and no alert is raised, resolved or suppressed because of it.
- Because a model call is an ordinary message on a queue, the console monitors it like any other: latency, failures, retries and dead letters all appear on the screens you already use.
Operator commands are permission-checked in the console, not just over MCP
If your host registers an authorizer, the console now enforces capabilities on every operator command — dead-letter replay and discard, listener pause/restart/drain, projection and agent operations, scheduled-message edits. Previously several of these were checked for an AI agent over MCP and not for a person clicking the same button, which is the opposite of what the capability list implies.
⚠️ If you have real permissions wired, read the upgrade notes — an operator who could run one of these before may now be denied until the capability is assigned. Hosts with no authorizer registered are unaffected: the default allows everything.
RBAC: the MCP and relay gates agree, and MCP actions work on an enforced host (#1258)
On a console that registers a real ICritterWatchAuthorizer, every MCP action tool reported accepted: true and did nothing. The tool authorized the caller and then published a command carrying no principal; the relay authorizes again, is fail-closed on a missing principal, and denied — without ever asking your authorizer, so no grant could have helped. The refusal had nowhere to go, because the caller was not a websocket, so the only trace was a log line. Hosts that never registered an authorizer were unaffected, which is why this went unnoticed.
The caller's identity now travels with the command, in this process, so the relay decides about the person who actually asked.
⚠️ Two capability families now name a different resource — tenant and projection commands. If your authorizer decides using the resource argument, see Upgrade Notes before upgrading.
Run the console inside your own application
CritterWatch no longer has to be a separate deployment. Two calls mount it inside a host you already run, monitoring that host.
opts.AddCritterWatchEmbedded(postgresSource); // inside your own UseWolverine — an NpgsqlDataSource on Marten,
// a connection string on Polecat and Fisher
app.UseCritterWatchEmbedded(); // mounts the UI under /critterwatch- No new package and no new project. Embedded is a registration mode on the three packages that already exist —
CritterWatch(Marten/PostgreSQL),CritterWatch.SqlServer(Polecat) andCritterWatch.Sqlite(Fisher). - The console's own data lives in its own store, isolated from yours by construction: it needs string stream identity where your application almost certainly uses
Guid, and that is a per-store setting. - Your routes are untouched. The UI mounts under a sub-path rather than taking the route table, so a host route that returns 404 today still does.
- The console gets its own
wolverine_*envelope tables — one runtime, two message stores. They are shared with the host only when the host has setopts.Durability.MessageStorageSchemaName(⚠️ not the same-named property insideIntegrateWithWolverine(o => ...)). A maintenance routine that rebuilds only the main message store will not touch them — see Embedded for the loop that covers every store. - The host is monitored automatically. Embedded has no separate service to instrument, so the console registers the monitored half on the host's behalf; telemetry travels in-process.
WARNING
Single instance only. Embedded self-monitoring rides in-process local:// messaging, which cannot cross nodes — so N nodes means N consoles writing the same streams, and no partitioning topology fixes it. The console detects and reports this at runtime. Use the standalone console for a clustered deployment.
Your host also needs WolverineFx.RuntimeCompilation (or its own codegen write covering both): the console's handler chains are generated at runtime here, because the static registries the standalone package ships with skip the assembly scan your own handlers are found by.
Every screen says how old it is
A console watching 514 databases and ~2,100 tenants rendered a 37-day-old snapshot as current, and nothing on any screen said so. That is the worst failure this product can have — an operator looking at a calm dashboard and concluding the fleet is calm — and most of this section is that gap.
- A console that has stopped applying telemetry now says so. Its self-ingest health check was attached to transport listeners only, and the queue where telemetry is applied is a local one, so the check never saw its failures: in a reproduction it answered
healthyover 849 constraint violations in 30 minutes. It now reportsdegradedwithin seconds, naming the exception. - Liveness alerts stand down against a frozen snapshot instead of measuring the console's own ingest gap as agent silence. One field console raised 25 Critical "agent down" alerts against agents last observed five weeks earlier. Alerts that read a snapshot's content still run, and a console with no ingest stamp is not treated as frozen, so upgrading does not silence anything on a healthy fleet.
- An observation age rides the rows, tiles and summaries that previously showed a value with no indication of when it was true. The one freshness indicator the console had measured the browser's own connection, which is green precisely when a stalled console looks healthiest.
- The landing page no longer states "No dead letter messages anywhere" over 15,972 of them for the first couple of seconds of every load, and the Operations tiles no longer claim "no events appended yet" over 2.66M events shown directly below them. Both were guards wired to a snapshot arrived rather than the counts arrived.
- Two ways a dead projection shard rendered as calm are closed. "Idle" conflated awaiting its first event with no agent is running it, and a frozen high-water mark subtracted to a healthy gap. Relatedly, 7,775 never-advanced shards no longer raise a permanent unmeasurable-lag warning that the same console calls Idle.
- The Timeline carries fleet events. It was 95.7% one flapping self-check — 180 of 188 entries over a month were the console watching itself, and no fleet event reached it at all.
Operating a fleet of this size
- The Operations route stops walking every progression row on mount. It fetched all 12,814 — 6.6 MB in 26 requests, 23 seconds of blocked main thread — to render a 25-row grid that reads none of them. Its header tiles are computed server-side now.
- The shard grid sorts and filters by health and lag, and the health chips above it are filters rather than labels. 467 pages with no sortable column is not a grid you can use.
- Total backlog is a real number. Summing each shard's gap would have been wrong by 3.3x, because the high-water mark is per (database, tenant) — shards on one tenant are not independent measurements. The roll-up accounts for that.
- A blue/green projection rollout is readable. The projection version is a dimension the server can group by rather than a substring; a shard row names the fleet that owns it; and each row carries a lag delta, so a rate and an ETA are measurable rather than eyeballed.
- Per-tenant alerting works above 25 tenants. Two independent ceilings used to collapse 2,173 tenants into a single roll-up whose membership was a top-five list inside a sentence. Tenants now ride the alert as data, a named few can keep their own alert above the ceiling, and one tenant can be silenced without going blind to its projection.
- An acknowledged alert is acknowledged everywhere, rather than staying acknowledged only in the tab that did it until every other console reloads.
- Several MCP tools reported a clean fleet from an empty source.
get_backlog_stateandlist_backlog_hotspotsread persistence counts that were{}while dead-letter summarize found 15,993 on the same console;list_degraded_surfacesreported every endpoint as degraded by excluding a status the endpoint vocabulary never emits;get_projection_lagreturned 11,656 unpaged rows with no tenant, store, gap or verdict. - "All stores" lists every store. An empty store identifier resolved to a single service, so the Streams tab could report six streams from an ancillary store as the whole fleet.
- The dev fleet boots from a Windows checkout. Without a
.gitattributesthe tenant-database init script was checked out CRLF and the Postgres container exited 127.
A recurring schedule says when it last fired
Following the Schedules that repeat section above, and worth calling out because it changes what the Schedule Explorer's numbers mean.
- Every runtime reading carries the instant it was true. A schedule's next-fire time and paused flag ride the capability announcement, which is only rewritten when the schedule's definition changes — so those readings were frozen at registration time and could sit hours stale with nothing saying so. Each service now reports its live schedule state on its own telemetry cadence, and the row says
as of 40s agofor a live reading orannounced 6h agofor one that is not. They are different claims and the screen words them differently. - "Is it still firing?" is answerable. A new Last handled column shows when a monitored node last actually ran an occurrence — an observed run, not the schedule's own bookkeeping. The two diverge exactly where it matters: a schedule whose message has no subscribers keeps a moving next-fire time and never handles anything. An empty column means not observed — an occurrence handled by a different service is invisible to it — and never idle.
- A pause you performed survives a reload. The console keeps its own record of every pause or resume a service verified, so the row no longer reverts to Running until the next announcement.
Check for updates
The About widget tells you when a newer CritterWatch is on NuGet, and links the release notes for the version you are running.
Fixed
"All stores" showed one store on a multi-store service. The Event Store Explorer's Configuration tab read the first event store in the service's capability document, which is host-registration order and not the main store — on a 514-database service that meant reporting three event types and one projection from an ancillary authorization store, with nothing on screen naming what had been read. The default is now the store the capability document identifies as the main one, and the picker displays its name. (The matching server-side read was fixed earlier in this release; this is the console half of the same assumption.)
An Event Store Explorer read that never came back left a spinner, not an error. These reads travel over the console's WebSocket, so a dropped answer leaves nothing behind — no failed request in the browser's network log, no error, no partial-result marker — and a reader could not tell a slow store from a lost response from a store with no streams. Every explorer read now has a 45-second deadline and reports a retryable failure naming what to narrow the read to, and the Refresh button no longer disables itself while a request is in flight.
The rollout panel could not always say whether a rollout was progressing, and said it anyway. Where every shard of a projection is unmeasurable, each version rendered a full row of blanks; the panel now collapses those to one line per projection, or to a single line for the whole panel when nothing on the service can be measured, and stamps the header with the snapshot's age when that snapshot is stale.
A projection command refused by a monitored service emitted no acknowledgement at all, so the console showed the action pending forever with the reason visible only in that service's log. Refusals now complete as refusals. Separately, Restart on a daemon-paused shard was a no-op that reported success — it lifted an operator pause and nothing else, so a shard the daemon had stopped stayed stopped while the console said the command had worked.
A pause no longer hides from the attention list. Paused shards were excluded from "needing attention", so the shard behind one August incident counted as healthy; a pause now needs attention unless an operator is the one who asked for it.
Ingest could latch under load. Per-service partitioning was not engaging on the highest-volume lane, so eighteen parallel upserts deadlocked and the console's shard view fell minutes behind. A retention trim deadlocked for a related reason — its ordering was not a total one, and every service trimmed the same window concurrently.
Ancillary event stores errored on every progression read and write, because they kept a bare progression table without the columns the console reads. CritterWatch now provisions that shape, and switches extended progression tracking on for every registered store through the store's own supported setting rather than inferring it from the container — see Upgrade Notes. A store that still reports the setting as off is now logged by name, which is a different failure from a read that returned nothing.
On a database-per-tenant PostgreSQL store the projection registry is readable now (marten#5382), so orphan detection works on the store shape that needs it most. The console reads that answer as a registry — every registered shard, no position — never as progression, so it cannot fork a shard onto an unattributed row.
Five field-reported defects carried over from the unreleased 1.0.2, all affecting fleets using something other than the default broker wiring:
- The console failed to boot on a plain broker listener. A console wired the way the docs and samples show it died at startup on Wolverine 6.24+ with
InvalidListenerConfigurationException: PartitionProcessingByGroupId() was configured on an Inline endpoint, naming a call the operator never made. CritterWatch applies that partitioning to every listener — it is the per-service single-writer guarantee — andInlineis simply the broker default. It now sets a mode that carries the partitioning, and leaves any listener you configured yourself alone. No configuration change is needed; if you addedProcessInParallelWithNativeAcks()as a workaround it remains correct and can stay. - Database-queue fleets silently stranded all telemetry. The console forced its Wolverine transport schema, so a
transportSchemapassed by an operator was overridden at host build. Services wrote telemetry, the console drained a different same-named table, the services endpoint returned[], and neither side logged anything. YourtransportSchemawins again — see Upgrade Notes. - Telemetry over the HTTP transport was rejected with
415. The telemetry route had been switched to a buffered mode, which makes an HTTP endpoint post a batch, and only the batch endpoint accepts those. The documented/_wolverine/invoketarget is correct again; if you moved to/_wolverine/batch/critterwatchas a workaround, that also still works. - The handshake log line named the wrong destination — a service reported sending its initial telemetry to its own control queue rather than to the telemetry URI, as did the publish-failure line, which is the one you read while actually debugging.
- An all-SQL-Server or all-PostgreSQL database-queue console served
500on every request (wolverine#4130). It booted and listened cleanly, then failed every request with… was NullMessageStore, because the outboxed session factory captured the Main message store before store reconciliation had assigned it. Fixed upstream and picked up by this release's pin. Store roles were correct throughout, which is why this presented as a packaging fault rather than a startup-ordering one; if you worked around it by registering the transport as ancillary, that remains a good configuration.
Supported stack
| .NET | 9.0 and 10.0 |
| Wolverine | 6.40.0 |
| JasperFx | 2.74.0 |
| Marten (PostgreSQL) | 9.38.0 |
| Polecat (SQL Server) | 5.30.0 |
| Fisher (SQLite) | 1.12.0 |
| Weasel | 9.32.0 |
| Vue | 3.5 |
Upgrading
Read Upgrade Notes before moving a running console to 1.1 — this release changes the progression-table shape on monitored stores, and finishes the transport-schema change begun in 1.0.2.
1.0.1
Released 20 August 2026 — nine fixes from the first production deployments of 1.0.0, most of them cases where the console reported a confident number it had no basis for.
Fixed
No alert had ever resolved. 33 Critical alerts sat stuck for 15–17 days at version 0 with resolvedAt null, while the monitored application was provably caught up — the console could not append to its own streams, so the resolution half of the alert lifecycle never ran. The single most serious defect in 1.0.0: alerting that can raise but never clear is worse than no alerting, because the board stops being read.
Counts of the same thing disagreed with each other. The alert bell said 28, the Health tab badge said 29, and the API returned 33 — all active, all Critical, none acknowledged. Separately, the Operations tab rendered "Dead Letters 0" in green beside a sidebar badge reading 999+ — two different counts under one label on one screen.
"Max Lag 0 · all projections caught up" was computed from a gap the console could not compute. An uncomputable gap is not a zero one, and reporting it as zero turns a blind spot into a reassurance.
An active alert carried no evidence it was still true — 33 Criticals rendered identically whether the claim was 40 seconds or 17 days old. Alerts now say how old their claim is.
The self-monitoring failure counter could not see the failures it exists for.selfMonitoring.recentFailureCount read 0 through 214 hard message failures.
500 blank Type cells, silently. CwTypeCell was used in six files without being imported; Vue emits an unknown <cwtypecell> element rather than failing, so /raw and the Metrics tab rendered blank columns with nothing in the console to say why.
Page navigation blocked for 8–18 seconds at fleet scale. projectionsForService cost ~1.7 s per call, was not memoized, and was called 3–4× per navigation; /projections blocked 17.8 s, /events 9.8 s, /documents 8.2 s. Memoized per service, the same treatment that fixed the connect path in 1.0.0.
Eleven Element Plus deprecation warnings across the route sweep (el-pagination small, el-radio label), which break outright at Element Plus 3.0.
Upgrading
No configuration changes are required from 1.0.0.
1.0.0
Released 19 August 2026 — the first stable release of CritterWatch.
CritterWatch is a monitoring and production-management console for Critter Stack systems. It gives you — and your AI agents — a live view of every Wolverine service, Marten and Polecat event store, and the controls to act on what you find.
1.0 marks the point where the API surface, the package lineup, and the storage model are stable and supported. Everything below has been through the full release gate on all three storage flavours.
Install
Pick the package matching the database you want CritterWatch's own console storage to use. This is independent of what your monitored services run — a console on SQLite can happily monitor a fleet of PostgreSQL-backed services.
| Package | Console storage | When to reach for it |
|---|---|---|
CritterWatch | Marten / PostgreSQL | The default. Most Critter Stack shops already run Postgres |
CritterWatch.SqlServer | Polecat / SQL Server | Shops standardised on SQL Server |
CritterWatch.Sqlite | Fisher / SQLite | Local development, demos, and single-node deployments — needs no database server at all |
dotnet add package CritterWatchThen see the Quick Start. Monitored services add Wolverine.CritterWatch, which is store-agnostic — it depends only on the JasperFx.Events abstractions, so a service on RavenDB, EF Core, or no event store at all can be monitored without pulling Marten in.
What is in 1.0
Understand your system, not just watch it. CritterWatch is as much a development tool as a production one. It renders live event models, your real store schema and DDL, the generated handler and projection source Wolverine actually runs, and full HTTP endpoint chains — the ground truth rather than your mental model of it.
The Projection Stepper replays a projection over a real slice of your event store and walks you through the before/after state row by row, with a JSON diff at each step. It is the fastest available answer to "why is this aggregate shaped like that at version N?"
Real-time monitoring of service health, nodes, agents, throughput and error rates, pushed over SignalR as events happen rather than polled.
Dead letter queue management across every service from one view — query, filter, replay, edit-and-replay, discard, in batches, with exception grouping and CSV export.
Projection operations — track async projection lag, spot stalls, and pause, restart, rebuild or rewind from the console with no code change and no service restart.
Event-sourced alerting with a complete lifecycle — raised, elevated, reduced, resolved, cleared — stored as immutable events, so the alert history is auditable rather than a current-state table that forgets.
Multi-tenancy — add, disable and remove tenants at runtime, with per-tenant metrics, DLQ filtering and projection health.
An MCP server for AI agents. Point Claude or any MCP client at one endpoint and it can query every monitored service — spans by saga or stream, projection lag, dead letters, alerts, metrics — and take action, tenant-scoped and RBAC-gated.
Licensing
CritterWatch is free to use in read-only monitoring mode. Connect as many services as you like and watch your whole Critter Stack with no license key.
A commercial license unlocks every administrative action — DLQ replay, projection rebuilds, listener and node control, tenant management — and the entire MCP server.
Supported stack
1.0.0 is built and gated against:
| .NET | 9.0 and 10.0 |
| Wolverine | 6.29.1 |
| Marten | 9.28.0 |
| Polecat | 5.19.0 |
| Fisher | 1.0.0 |
| JasperFx / JasperFx.Events | 2.52.1 |
Upgrading from a release candidate
No configuration changes are required from 1.0.0-rc.10.
⚠️ If you are coming from a build older than rc.10 and ran an embedded console, note that CritterWatch.Embedded was removed before 1.0 — the console now runs as its own host. Its storage collapsed to a single store at the same time: console data lives in the critterwatch schema and Wolverine durability in critterwatch_wolverine. No table moved, but an upgraded database may still carry orphaned critterwatch.wolverine_* tables from the old two-store layout. Nothing reads them and they can be dropped.
Getting help
1.0 closed 282 issues across ten release candidates, beginning 27 July 2026.
Questions, bug reports and feature requests are welcome in the Critter Stack Discord, and licensed customers can reach us through the support address on their license.
