Services / IT administration
OperationsDay-2 with an owner, a change window, and an audit trail
We take over or support operations: hosts, virtualization, identity, copies, and on-call. Proxmox or another platform has a map. Grafana shows health. Backup has a restore test. A change has a window, not chance.
One systems map, one escalation channel
Administration without observability reacts after the fact. Backup without restore is an assumption. We put the platform console, cluster, alerts, and copies into one on-call.
- Patch management covers hosts, hypervisors, and guests, with order depending on HA.
- Identity and MFA are part of the retainer, not a separate project once a year.
A datacenter you can see in one console
Proxmox VE on a cluster: nodes, VMs, containers, storage, and HA. Patch and live migration of a running machine are on the calendar. Root access is not a daily tool.
Inventory
Machines, CTs, network, and storage with an owner. Golden images instead of a manual installer on every VM.
Change windows
Node and guest updates with PDB or HA order. Rollback is in the plan, not in hope.
Access
Named accounts, 2FA, a separate jump. A log of every login and task in the cluster.
Capacity
CPU, RAM, and storage with a threshold. Expansion before a guest starts swapping.
Containers under the same care as VMs
When some services run on a cluster, administration covers nodes, certificates, storage class, and upgrade. We do not leave Kubernetes as an island outside on-call.
Nodes
OS, kubelet, and runtime patch in a change window. Drain and cordon, not a reboot at noon.
Add-ons
Ingress, DNS, storage, cert-manager. Versions and configuration backup are written down.
Observation
Pod and node health in Grafana, not only in an ad hoc dashboard.
Boundary
RBAC and NetworkPolicy. Administration does not mean permanent cluster-admin.
An alert before the user calls
Host, VM, cluster, and service have a dashboard. Disk, certificate, backup, and queue have a threshold. On-call gets a priority and a runbook, not a Nagios dump without context.
Scope
Linux, Windows where it is present, database, queue, HTTP. One silence list for maintenance.
Priority
P1 is loss of service. P3 is an approaching certificate. Not everything wakes you at three.
Correlation
Host and application side by side. A storage failure does not look like a mysterious API timeout.
Report
Week: how many alerts, how many change windows, what stayed open. Material for a review with you.
A copy you can restore
Proxmox Backup Server or a second path outside the site. Retention, immutability, and a restore test are in the contract. A failed job is an alert, not a silent email.
Scope
VMs, CTs, databases, and configuration. Vault secrets outside the same bucket as application data.
RPO and RTO
Written per system. Restore of a single file and of a whole machine are practiced.
Separate path
A second copy on other accounts. Ransomware in production does not delete the only copy.
People
Who restores, in what order, with which DNS. A playbook available offline.
From takeover to a weekly rhythm
First inventory and access. Then a monitoring and copy baseline. Finally change windows, IAM, and a report.
- Takeover Map of hosts, accounts, copies, who has the key, which SLA.
- Baseline Monitoring, backup, first change windows, removal of unnecessary access.
- Rhythm Patch, alert review, restore test, change request.
- On-call Escalation, report, material for a change audit.
We will discuss the systems map, on-call, and change windows
On that basis we will prepare a scope of administrative care.
Contact us