Dashboards nobody owns
Several near-identical boards built by different teams for the same service. Nobody knows which is authoritative, so during an incident people check two and trust neither.
Platform expertise / Grafana
It is the most common thing we are shown: hundreds of panels, several near-duplicates of each other, and an on-call engineer who still opens a terminal first. A dashboard earns its place by answering a question somebody asks during an incident. Most do not, and the useful work starts by finding out which.
Grafana is third-party software selected and licensed by the client. Aevis provides advisory, engineering and operational services around the clientโs deployment, whether that is self-hosted open-source Grafana, Grafana Cloud or Grafana Enterprise.
Observability layer
Instrumented ยท correlated ยท alerted ยท answeredPlatform fit
Grafana is easy to add a panel to, which is its great strength and the source of the problem. Dashboards accumulate faster than anybody removes them, alerts are added after each incident and never retired, and the signal that would have explained tonight is not on any of them.
Several near-identical boards built by different teams for the same service. Nobody knows which is authoritative, so during an incident people check two and trust neither.
Thresholds set on component state rather than on customer symptom. On-call has learned which ones to ignore, which is the point at which the alerting system stopped working.
The metric shows a spike and the log that explains it lives in a different tool with a different time filter and a different idea of what the service is called.
Our role is to make the deployment answer the questions your on-call actually asks โ not to sell an Aevis software product.
Product landscape
We shape the engagement around the components and edition your organisation has selected. Feature availability, support and cost model differ significantly between open-source, Cloud and Enterprise.
The visualisation and exploration layer, and the alerting engine that increasingly does the important work.
The metrics store, its cardinality behaviour, and the recording rules that decide whether queries stay affordable.
Log aggregation designed around labels rather than full-text indexing โ cheap if the label design is right, painful if it is not.
Distributed tracing and the sampling decisions that determine whether the trace you need is the one you kept.
Routing, escalation, silences and the schedule โ the part of observability that touches people directly.
Probing from outside the estate, and load testing that exercises the paths real users take.
Aevis capabilities
Engage us for a focused intervention or an end-to-end programme. We work within your licensing, hosting and data-protection constraints.
Which boards are opened during incidents, which alerts have ever led to an action, and which signals nobody is collecting.
Alerting on what a user experiences rather than on what a component is doing, which is what makes a page worth waking somebody for.
Getting the right signals out of applications and infrastructure, labelled consistently enough to correlate.
A small set of boards designed around the questions asked in an incident, and the removal of the rest.
Running the stack itself: capacity, cardinality, retention and the cost behaviour that follows from all three.
Running the platform, the alert estate and the on-call configuration, or standing behind a team that does.
AI and analytics in observability
Anomaly detection and assisted correlation are genuinely useful on telemetry, which is high-volume and numeric. They are also the capabilities most often oversold, so the line below states where the output stops being evidence and starts being a decision.
Baselining per service and per season, so an alert can fire on a departure from normal rather than on a static threshold somebody guessed.
Collapsing a storm of related alerts into one incident with a probable common cause, ranked for a human to confirm.
Drafting PromQL or LogQL and summarising an incident timeline, with the query and its result shown rather than hidden.
What stays human โ without exception
No AI declares an incident resolved, suppresses an alert or authorises a remediation action. Cause, impact and the decision to act are engineering judgements made under your incident process, and remain the accountable decision of the person who made them. A ranked probable cause is a starting point for investigation, and a page that treats it as a conclusion has produced a faster wrong answer.
How value is measured
Entitlement and data
Which analytics and assistive capabilities are available depends on whether the client runs open-source Grafana, Grafana Cloud or Grafana Enterprise, and on the version deployed. Telemetry is processed for the agreed operational purpose only, under the clientโs data-protection terms.
Connected architecture
Metrics, logs and traces answer different questions at very different costs. Collecting all three at full fidelity for everything is affordable for nobody, and the design work is deciding what each layer is for.
Application and infrastructure signals, their labels, and the naming that decides whether anything correlates.
Agents, scraping, sampling, cardinality and the retention tier each signal class earns.
Service boards, SLO alerting, notification policy and the on-call rota it feeds.
Service management, the SIEM, the CMDB and whatever else needs to agree with this on what a service is called.
Architecture boundaryAvailable features, high-availability options, retention behaviour and support differ materially between self-hosted open-source Grafana, Grafana Cloud and Grafana Enterprise. We confirm which the client runs, and its version, before committing to a design.
Delivery model
The fastest way to find out what an observability estate is missing is to walk through the last three incidents and note every question that took more than a minute to answer.
Walk recent incidents, measure dashboard and alert usage, and list the signals that were needed and absent.
Signal gap list and alert-to-action baselineDefine service levels, what deserves a page, what deserves a ticket, and who owns each serviceโs observability.
SLOs, alert policy and named ownersInstrumentation, label design, a small set of provisioned dashboards and symptom-based alert rules.
Dashboards as code and symptom alertingRun the new alerting in parallel with the old for a full on-call cycle before anything is retired.
A cycle of evidence before cutoverRun the platform, the alert estate and the cost position, with post-incident findings fed back in.
Governed operating cycleRetire boards nobody opens and alerts nobody acts on; close the signal gaps each incident reveals.
Smaller estate, higher alert precisionUse cases
Each of these is a normal starting point rather than a programme. We map the adjacent dependencies so a local fix does not create a hidden failure elsewhere.
Thresholds on component state rather than customer symptom. Precision matters more than coverage once people are filtering by instinct.
Usage data usually shows a small handful are opened at all. The rest can go, once each survivor has an owner.
A label with unbounded values multiplied the series count. It is a design problem with a design fix rather than a capacity purchase.
Different naming in each tool. Consistent service identity does more for time-to-locate than any new dashboard.
Moving from several tools, with parallel running until the boards and alerts that matter behave identically.
Service levels defined with the owning team, so health is a stated target rather than an aggregate of green panels.
Engagement shapes
Which one fits is usually a question about where accountability should sit rather than about budget.
Best forAlerts nobody trusts
A bounded assessment across recent incidents, dashboard usage and alert-to-action data, ending in a prioritised gap list with an owner against each item.
Best forInstrumentation, SLOs or consolidation
Defined scope with acceptance criteria โ instrumentation, alert redesign, dashboards as code or a migration โ handed over with the design documented.
Best forNo standing platform team
Aevis operates the stack, the alert estate and the cost position to an agreed cadence, with the accountability boundary set out in the service agreement.
Best forA team that should own this
We work alongside your engineers and hand over deliberately, with dashboards provisioned as code and train-the-trainer where the capability should stay with you.
Designed outcomes
Baselines and targets are agreed per engagement. We do not import a vendor benchmark into your estate and call it a business case.
Alerts leading to an action, as a share of alerts fired.
How long it takes to place a fault at a service and a component.
Boards opened during incidents, against boards maintained.
Series, log volume and trace retention against the questions they serve.
No provider can guarantee availability or that an incident will be detected before a customer notices. What is contracted is the engineering, the operation and the improvement practice within an agreed scope; the organisation retains its service commitments and its risk decisions.
Governance
Observability estates grow in one direction unless something removes from them. These are the standing controls that supply the other direction.
A dashboard without a named owner is a candidate for deletion at the next review, which is the only mechanism that reliably shrinks the estate.
Dashboards and alert rules live in version control, so a change has an author, a reason and a way back.
Alerts are reviewed against whether they led to an action; the ones that never did are retired rather than silenced.
Label design is reviewed before a source is onboarded, because an unbounded label is a cost incident with a delay on it.
Why Aevis
We approach Grafana as a system somebody is on call behind. The work is designed to survive handover, a noisy night and a change of team.
The gap list comes from walking real incidents and noting the questions that took too long, not from a maturity model. It produces a shorter and more specific piece of work than an assessment framework does.
The recommendation is usually a deletion, and it is worth less revenue to us than building more. A board nobody opens is maintenance cost with no return.
The people designing your alerts have been woken by bad ones. What is worth a page at three in the morning is argued about from experience rather than from a threshold template.
We instrument with OpenTelemetry and provision as code, so the work has value even if you later move platform. A design that only functions on one vendor is a design we would have to defend rather than justify.
Relationship clarityAevis does not claim ownership of Grafana products and this page does not state or imply a certified partnership. Product names and trademarks belong to their respective owners.
Testimonials
Each testimonial is tied to the service it refers to, so service pages can draw the relevant one automatically.
The change we noticed first was not technical. It was that there was finally one person to call, and that person already knew the history of the problem.
They rebuilt the service catalogue around how our teams actually work rather than how the platform was shipped. Adoption stopped being an argument.
We had the security tooling before Aevis arrived. What we did not have was anybody turning what it produced into decisions.
Frequently asked questions
The useful answers depend on your deployment and edition. These are the principles we use before an assessment establishes the exact scope.
This page makes no partnership claim. Aevis provides advisory, engineering and operational services around a deployment the client licenses or self-hosts. Where a formal partner relationship is relevant to a procurement, ask us and we will answer it precisely rather than by implication.
It changes a good deal, and it is the first thing we establish. Features, high-availability options, retention behaviour and support differ materially between open-source, Cloud and Enterprise. A recommendation that assumes Enterprise capability on an OSS deployment is simply wrong, so we confirm the edition and version before designing anything.
They answer different questions and the overlap is worth deciding deliberately. Metrics-first questions about latency, saturation and error rate are usually cheaper and faster in this stack; questions about log content, investigation and retention obligation typically belong in the SIEM. What matters is that the same data is not paid for twice, and that both agree on what a service is called.
Some of them, and only with usage data behind the recommendation and an owner asked first. Nothing is removed on our say-so. But an estate where nobody can identify the authoritative board for a service is one where an incident starts with a search, and the fix for that is subtraction rather than another board.
Usually not first. Most of the growth we are shown is cardinality โ a label with unbounded values multiplying the series count โ or retention set by default rather than against a question. Both are design problems with design fixes, and both travel with you if you change vendor.
That is the co-managed shape, and this platform is particularly well suited to it. Dashboards and alert rules are provisioned as code so they are readable and reversible, the design is documented, and train-the-trainer is available through the Corporate Training practice.
Grafana enquiry
Tell us about the most recent incident that took too long to diagnose, and which questions during it were slow to answer. That is a more useful brief than a list of the tools you run.