Kaizen Official Group

Services

Every service has a “what is not included” list. It removes more objections than the list of what is included, and it protects both of us from having different ideas of the word “done”.

Areas

Eight areas of work

Monitoring, logs and alerting

Discover failures before your customers do.

What we can do

  • Prometheus, Grafana, Loki, Alertmanager or another suitable stack
  • Availability, CPU, memory, disk, TLS certificates and containers
  • Infrastructure and application logs
  • Dashboards showing the state of the system
  • Telegram alerts that carry the reason for the failure
  • Event grouping and noise reduction
  • A controlled failure to prove the signal actually arrives
  • Operational documentation

Not included

  • Being on call for your alerts around the clock
  • Rewriting your application to make metrics convenient

Result: a working observability setup, verified notifications and a clear picture of your infrastructure.

Timeline and price — after we review the task

Discuss monitoring

Linux server and access audit

Find infrastructure weaknesses before they cause an incident.

What we can do

  • Services and ports reachable from the internet
  • Users, SSH keys and shared accounts
  • Permissions, updates, firewall and SSH configuration
  • Unnecessary services and unsafe settings
  • Closing off access that is not needed
  • Safe rotation of keys and credentials
  • A report with priorities and recommendations

Not included

  • Penetration testing and intrusion attempts
  • Legal assessment of regulatory compliance

This is an engineering audit of configuration, not a penetration test and not a compliance assessment.

Result: a map of risks, unnecessary access closed and a documented improvement plan.

Timeline and price — after we review the task

Request a server audit

Incident diagnosis and stabilisation

Fix the cause, not just restart the service.

What we can do

  • Metrics, logs and system state examined
  • Disk, memory, CPU, processes and network checked
  • Containers and background jobs reviewed
  • Timeouts, restarts and resource exhaustion diagnosed
  • Problem dependencies and points of failure identified
  • A reversible plan of changes
  • The agreed fix applied and the result verified
  • Monitoring, alerts and a recovery runbook

Not included

  • Changes to application business logic
  • Large migrations — quoted separately

Result: a confirmed cause, a verified fix and a lower risk of the failure repeating.

Timeline and price — after we review the task

Investigate an incident

Docker and reproducible environments

Run your application consistently on every prepared server.

What we can do

  • Dockerfile and Docker Compose
  • Application, configuration, data and secrets separated
  • Volumes, networks and environment variables
  • Healthchecks, restart policies and resource limits
  • Versioned image builds
  • Moving an existing application into containers
  • Instructions for start, update and rollback

Not included

  • Moving to Kubernetes without a confirmed need
  • Building application features

Result: a reproducible environment, a predictable start and a documented update procedure.

Timeline and price — after we review the task

Containerise a project

Automated deployment and CI/CD

Release a website, CRM, bot or service without manual file copying.

What we can do

  • Build pipeline and automated checks
  • A versioned artefact or Docker image
  • Storage for images and packages
  • Secrets moved out of the repository
  • Test, stage and production environments
  • Automated deployment
  • Database migrations in a controlled order
  • Healthcheck verification after release
  • Rollback to the previous version
  • Notifications about the release outcome
  • Release and recovery documentation
  • GitHub Actions, GitLab CI/CD, TeamCity, Jenkins, Docker — whichever fits

Not included

  • Building a new CRM from scratch — a separate product project
  • Replacing CI/CD with Terraform or Ansible: they have another role

Suitable for websites, shops, off-the-shelf CRM and internal systems, Telegram bots, APIs, microservices and Docker applications on a single VPS or several servers. Deploying an existing CRM is included. Terraform provisions reproducible infrastructure and Ansible configures servers — both complement CI/CD rather than replace it.

Result: a repeatable, verifiable release tied to a code version, with a rollback you can rely on.

Timeline and price — after we review the task

Automate deployment

Backups and verified recovery

A backup is useful only when it can be restored.

What we can do

  • Deciding which data and configuration must be copied
  • Database and file backups
  • Copies stored separately from production
  • Retention and rotation
  • Monitoring the success and age of the latest copy
  • Notifications when a backup fails
  • A test restore
  • RPO and RTO agreed with you
  • An emergency recovery runbook

Not included

  • Recovering data that no longer exists anywhere
  • Storage costs carried by the contractor

Result: a verified backup process and a documented recovery scenario.

Timeline and price — after we review the task

Verify backups

Kafka and reliable message flows

Deploy Kafka from scratch or stabilise an existing system.

What we can do

  • Kafka on a VPS, dedicated servers or in Docker
  • A single broker or a small cluster
  • KRaft configuration
  • Topics created and configured
  • Partitions, replication factor and retention chosen for the load
  • Network access, authentication and secrets
  • Producers and consumers verified from the infrastructure side
  • Consumer lag, offsets and stuck consumers diagnosed
  • Producer and consumer errors investigated
  • Metrics, Grafana and alerts
  • Basic failure and recovery scenarios tested
  • Operational documentation

Not included

  • Changes to producer and consumer application code
  • Large high-load platforms and complex migrations — quoted separately after a load analysis

Result: a working, observable Kafka setup, or an existing system diagnosed and stabilised.

Timeline and price — after we review the task

Discuss Kafka

Bots and operational automation

Remove repetitive manual work from daily operations.

What we can do

  • A Telegram bot for a specific operational task
  • Reports and notifications
  • Regular checks and repetitive operations automated
  • The bot connected to an API, database or monitoring
  • Safe administrative commands
  • Logging and error tracking
  • The solution deployed on your server
  • Source code, documentation and access stay with you

Not included

  • Open-ended product development
  • Unknown legacy without a review first — quoted separately

Result: a working tool that cuts manual work and makes operations repeatable and controlled.

Timeline and price — after we review the task

Automate a task

After delivery

Ongoing support

Keeping an eye on what was delivered, checking alerts and backups, reviewing errors, agreed updates, a periodic report on state and risks, and a limited amount of engineering work.

Scope, availability hours and response time are agreed separately. Round-the-clock duty, instant response and unlimited incident work are not promised.

Discuss ongoing support

How it works

Four steps, no surprises

  1. Review

    I look at what you have and name the problems plainly.

  2. Plan

    In writing: what I do, in what order, and what you get at the end.

  3. Work

    Every step is planned to be reversible, with the rollback prepared before anything changes.

  4. Report

    What was done and how it was verified. The infrastructure stays yours — access, documentation, all of it.

Not sure which one you need? That is normal — we work it out together when we review the task.