NetVoice · custom application development

Custom applications built by a team of AI agents. A human makes the call.

Agents write, review and test code on GitHub, Azure and Cloudflare. You say what you need and approve the result.

Built on
  • Anthropic
  • GitHub
  • Microsoft Azure
  • Cloudflare
  • 5× faster delivery of an application internal web application
  • 2 AI agents review every change
  • 100% of changes are released by a human
Request an application
How it works

Five steps from brief to finished application

  1. Human

    You describe

    In your own words, you describe what the application should do.

  2. AI agent

    AI agents build

    Specialized AI agents split the work and write the code.

  3. AI agent

    Another AI agent reviews

    Every change is reviewed by an AI agent that did not write it.

  4. AI agent

    An AI agent tests

    Automated tests try out the application in a browser and on mobile.

  5. Human

    You approve

    Only what you sign off goes live.

Every step can be traced on GitHub, and deployments run on Azure or Cloudflare.

Night mode

AI agents keep working at night, when nobody is waiting. In the morning, finished and tested changes are ready for your approval.

Example: fixing a bug step by step

Brief:

    Bug fix: “The form can't be submitted on mobile.”

    1. 09:02 · Human → Dispatcher · Message The form can't be submitted on mobile. Customers are complaining. A person describes the problem in their own words. Nothing more is needed.
    2. 09:04 · Dispatcher → Developer (frontend) · Issue Creating the issue “Form can't be submitted on mobile”. The frontend developer takes it, priority is high. The Dispatcher turns the report into a specific issue and decides who takes it and how urgent it is.
    3. 09:31 · Developer (frontend) → Reviewer · Change for review Done. On small screens the button was covered by another element. Sending it for review. The Developer works on its own copy of the code. It submits the change for review and cannot push it to production itself.
    4. 09:40 · Reviewer → Developer (frontend) · Code review The fix works, but there's no test that would catch the same bug next time. Sending it back. The review is done by a different agent than the author. It sent the change back because nothing prevented the bug from coming back.
    5. 09:52 · Developer (frontend) → Reviewer · Change for review The test is added and the automated checks passed. The author addresses the comment and resubmits the change.
    6. 09:55 · Reviewer → Tester · Code review Now it's fine. Approved from a code perspective, passing it to the Tester. The code review passed. But the change is not in production yet.
    7. 10:10 · Tester → DevOps · Test Tested in the browser on three mobile sizes. The form submits. The Tester tries the fix the way a customer would use it.
    8. 10:14 · DevOps → Human · Release The fix is running on a staging URL and awaits your approval. The staging version (preview) is separate from production; customers don't see it.
    9. 13:05 · Human → DevOps · Approval I've checked the preview, approved. This is the only step that releases the change to production. It is always done by a human.
    10. 13:07 · DevOps → Production · Release Deployed to production; the post-deployment check is fine. Deployment runs automatically only after approval and can be rolled back to the previous version.

    Times show an illustrative run; the example is a model.

    What you save

    Same hourly rate, up to 5× fewer hours

    We charge the same rate as a typical supplier (1,500 CZK per hour excl. VAT). But agents need far fewer hours.

    • App development 5×faster

      Internal web application

      For example, attendance tracking: sign-in, clocking in and out, monthly overviews, export and user management.

      Programmer at a supplier 60 working days (≈ 3 months)
      NetVoice with agents 12 working days (≈ 2.5 weeks)
      Supplier≈ 720,000 CZK NetVoice≈ 144,000 CZK

      70–80% cheaper you save ≈ 576,000 CZK

      Of which 2 working days of human work (brief, ongoing checks and approval).

    • User interface design 4×faster

      Design of 10 screens including mobile

      A designer draws in Figma; the agent designs the screens directly in code following the design system and prepares variants to choose from.

      UI designer (Figma) at a supplier 10 working days (≈ 2 weeks)
      NetVoice with agents 2.5 working days
      Supplier≈ 120,000 CZK NetVoice≈ 30,000 CZK

      70–80% cheaper you save ≈ 90,000 CZK

      Of which 4 hours of human work (choosing variants and giving feedback).

    • Operations and infrastructure 3.3×faster

      Complete Azure environment

      Network, servers, databases, backups, monitoring and access rights, all described as code so the environment can be recreated at any time.

      DevOps specialist at a supplier 10 working days (≈ 2 weeks)
      NetVoice with agents 3 working days
      Supplier≈ 120,000 CZK NetVoice≈ 36,000 CZK

      70% cheaper you save ≈ 84,000 CZK

      Of which 1 working day of human work (review of the design and security, approval).

    Comparison with your own employee

    When your own employee does the work, you pay their salary and contributions for the hours worked.

    Project Own employee NetVoice Difference
    Internal web application ≈ 334,000 CZK ≈ 144,000 CZK 56% cheaper
    Design of 10 screens including mobile ≈ 56,000 CZK ≈ 30,000 CZK 46% cheaper
    Complete Azure environment ≈ 56,000 CZK ≈ 36,000 CZK 35% cheaper

    Employee cost of 117,000 CZK/month = average salary of 87,153 CZK + 33.8% employer contributions (social security and health insurance that Czech employers pay on top of the gross salary). Employee hour = monthly cost / 168 hours = 696 CZK.

    Sources for rates and salaries: INITED Solutions: agency hourly rate 2,000 CZK (2024) · Agionet: design and programming 1,100–1,250 CZK/h · Zdeněk Skulínek: senior programmer 1,500–2,500+ CZK/h · Jooble: average salary in software development 87,153 CZK/month (Sep 23, 2026)

    References

    Applications we have built this way

    Personal and business data are blurred in the screenshots.

    • Doklady AI: received invoice in PDF next to the extracted data with a confidence score for each field

      Doklady AI

      An accounting platform that reads invoices from PDF and keeps bookkeeping and VAT in check.

      • Extracts invoice data and shows how confident it is.
      • Issues invoices with a QR payment code and repeats them the next month.
      • Exports documents to Czech accounting software (Pohoda, Money, Helios, ABRA).
      App tour: overview, received invoices, extracted data next to the PDF and an issued invoice with QR code
    • Docházka: monthly overview of hours worked, balance and overtime

      Docházka

      Attendance tracking for smaller companies, usable from a mobile phone too.

      • Clock in and out with a single button.
      • Overview of hours, overtime and vacation.
      • Monthly closing as the basis for payroll.
      App tour: month view, leave approval, team calendar and monthly closing
    • Energy Monitoring: plant overview with power demand, consumption and costs

      Energy Monitoring

      Continuous energy metering in a manufacturing plant with cost monitoring.

      • Warns before the contracted capacity is exceeded, before a penalty is due.
      • Reports leaks, outages and unusual consumption.
      • Sends reports to management and supporting data for audits.
      Short app tour
    • Rónin: valley with a river and blossoming cherry trees

      Rónin

      A 3D game for web browsers, Android and iOS, developed by a team of specialized AI agents.

      • In 5.5 days, 7 AI agents: 1,044 code edits, 254 approved changes and 82 closed tasks.
      • Agent cost so far ≈ 8,000 CZK; the same number of people costs 819,000 CZK a month.
      • Target for a playable version of the whole game: October 3–5, 2026; a human always approves.
      Gameplay: the valley by day, the camp at dusk and a boss fight
    • 3D viewer: overall view of a scanned facility

      3D scans

      A web viewer for 3D scans of buildings and facilities (LiDAR, drone) – viewing, measuring, tagging elements.

      • Fly through and walk around the point cloud right in the browser.
      • Measure distances and areas.
      • Place tags and notes right next to elements in the scan.
      Fly-through of a warehouse, a measurement and a tag with details
    Why it is safe

    Four safeguards on every change

    • Review by another agent

      Code written by one agent is always reviewed by a different agent.

    • Automated tests

      Tests verify every change before you ever see it.

    • A human approves

      Nothing goes live without your approval.

    • Every change can be undone

      Every release is recorded and can be rolled back one step.

    Your own agent team

    Strengthen your team with agents

    Agents can also work directly in your team. Your people brief and approve; the agents help them get the work done.

    Lower monthly costs ≈ 430,000 CZK model team of 6 people, calculation below
    Calculate for our company

    Savings by role

    • Programmer

      3 people

      ≈ 324,000 CZK / mo

      today ≈ 351,000 CZK agents ≈ 27,000 CZK per year ≈ 3,886,000 CZK

    • UI designer (Figma)

      1 person

      ≈ 114,000 CZK / mo

      today ≈ 117,000 CZK agents ≈ 3,000 CZK per year ≈ 1,366,000 CZK

    • Tester

      1 person

      ≈ 113,000 CZK / mo

      today ≈ 117,000 CZK agents ≈ 4,000 CZK per year ≈ 1,358,000 CZK

    • DevOps specialist

      1 person

      ≈ 116,000 CZK / mo

      today ≈ 117,000 CZK agents ≈ 1,000 CZK per year ≈ 1,386,000 CZK

    Use the slider to set the size of your team. For each person we count one agent of that role; a server is added for every 10 agents. Fine-tune everything in the calculator.

    Net saving per month after deducting the dispatcher, servers and 2 people in management

    ≈ 430,000 CZK/mo Net saving per year ≈ 5,157,000 CZK/yr

    Model values. Adjust the detailed assumptions in the calculator below.

    Traditional team
    Role People CZK / person / mo
    Programmer
    UI designer (Figma)
    Tester
    DevOps specialist
    Agents
    Role Count USD / agent / mo Helps
    Dispatcher shared
    Architect Programmer
    Developer Programmer
    Reviewer Programmer
    UI Designer UI designer (Figma)
    Tester Tester
    DevOps DevOps specialist
    Management
    Servers and exchange rates

    Monthly saving by role

    • Programmer ≈ 324,000 CZK
    • UI designer (Figma) ≈ 114,000 CZK
    • Tester ≈ 113,000 CZK
    • DevOps specialist ≈ 116,000 CZK
    • Shared costs ≈ −237,000 CZK
    Net saving per month ≈ 430,000 CZK/mo
    Net saving per year ≈ 5,157,000 CZK/yr
    Traditional team
    ≈ 702,000 CZK/mo
    AI team incl. people in management
    ≈ 272,000 CZK/mo
    Calculation details exact values in CZK per month and year

    Agent cost = count × price in USD × CNB exchange rate. Traditional team cost = number of people × cost per person. Saving per role = traditional team − the agents helping in that role.

    Role Traditional team / mo Helping agents Agents / mo Saving / mo Saving / yr
    Programmer 351,000 CZK 3 × 117,000 CZK Architect, 3× Developer, Reviewer 27,126 CZK 323,874 CZK 3,886,489 CZK
    UI designer (Figma) 117,000 CZK 1 × 117,000 CZK UI Designer 3,204 CZK 113,796 CZK 1,365,554 CZK
    Tester 117,000 CZK 1 × 117,000 CZK Tester 3,845 CZK 113,155 CZK 1,357,865 CZK
    DevOps specialist 117,000 CZK 1 × 117,000 CZK DevOps 1,495 CZK 115,505 CZK 1,386,058 CZK
    Shared costs – Dispatcher, servers, people in management 236,546 CZK −236,546 CZK −2,838,557 CZK
    Total 702,000 CZK 272,216 CZK 429,784 CZK 5,157,408 CZK

    Shared costs include agents not tied to a specific role (the dispatcher), servers, and the people who stay in management and approvals. They are deducted from the sum of savings across roles.

    Model assumptions

    The calculator works with model assumptions, not with a quote. Actual costs depend on the chosen model, the volume of work, context length and the current API price list. Prices per agent are indicative monthly averages. Exchange rates follow the Czech National Bank (CNB) rate list of September 25, 2026 (1 USD = 21.359 CZK, 1 EUR = 24.350 CZK) and are refreshed with every site build. The cost per person includes salary and employer contributions, but not licenses, equipment or administration. The comparison does not account for onboarding time or differences in output quality.

    FAQ

    What managers ask most often

    How much does it cost?

    You pay the same hourly rate as with a regular supplier, just for far fewer hours. You get the exact price upfront in our quote.

    Is our data safe?

    AI agents work with code and made-up sample data. Live data stays in your environment, accessible only to people you choose.

    Who is responsible for the result?

    NetVoice is responsible for the delivery. Every change is reviewed by another AI agent and released by a human.

    What if the AI makes a mistake?

    Reviews and automated tests catch most errors before they reach you. If an error does make it into production, we roll back to the previous version.

    How do we get started?

    Tell us what the application should do and we will arrange a short meeting. You then get a quote and usually a first preview within a few days.

    Who owns the finished application?

    The application and its source code belong to you. You can keep working with them even without us.

    Contact

    Need an application?

    Tell us what the application should do and who it is for: an internal system, a customer portal, an accounting integration or automation. We will estimate the price and timeline and suggest where to start.

    Send an email

    netvoice@netvoice.cz

    NetVoice s.r.o.
    Pobřežní 249/46, 186 00 Praha 8 – Karlín
    Company ID 27592570 · VAT ID CZ27592570
    For IT

    For IT specialists: how it works inside

    Architecture, agent communication on GitHub, servers and costs. Every tab can be shared as a link.

    Architecture

    Architecture

    Workflow from requirements to production

    Every step has a single owner. Bugs found by tests go back to the dispatcher as new tasks, not straight to the developer. Below the diagram is a model night run step by step, including how it looks on GitHub.

    The night run step by step

    A model run: a bug found at 1:10, fixed, reviewed and tested by 2:46. In the morning a human approves it and releases it to production.

    Agents react to labels and comments on GitHub, not to chat.

    Step 1 of 12: Evening: a human approves the night queue

    Tip: you can also switch steps with the ← → arrow keys or by clicking a node in the diagram.

    Step 1 of 12 · evening node: Human

    Evening: a human approves the night queue

    1. status: new
    2. status: in progress
    3. review
    4. testing
    5. awaiting approval
    What happens

    The product owner goes through the queue and uses the night label to mark what agents may work on overnight. Merging to production is left for the morning.

    On GitHub

    Night queue: cart and payments #248

    Open product-owner opened this issue

    Labels
    plannight
    Assignees
    agent-dispatcher

    product-owner added night labels · 18:29

    product-owner commented · 18:30

    @agent-dispatcher approving the night queue: cart, payments and the nightly e2e tests. Take fixes as far as a preview; leave the merge for the morning.

    Message between agents

    product-ownerhuman to agent-dispatcher

    via night labelcomment@mention

    “Approving the night queue: cart, payments and the nightly e2e tests. Take fixes as far as a preview; leave the merge for the morning.”

    Step 2 of 12 node: Tester

    The nightly test fails

    1. status: new
    2. status: in progress
    3. review
    4. testing
    5. awaiting approval
    What happens

    At 1:00 the e2e tests start automatically against main. Two discount cases fail and the tester reports it to the dispatcher.

    On GitHub

    nightly e2e · Failure
    schedule · 01:00 · main @ a41c9d7

    1. passed: checkout, npm ci 41 s
    2. passed: unit tests 209 of 209
    3. failed: playwright e2e 46 of 48

    e2e cart: 2 failures

    • passed: cart: checkout without discount
    • failed: 10% discount + 21% VAT off by 0.01 CZK
    • failed: 15% discount + 12% VAT off by 0.02 CZK

    agent-tester commented · 01:10 · in #248

    @agent-dispatcher nightly e2e on main (a41c9d7) failed in 2 of 48 cases: discount + VAT in the cart. The run report is attached.

    Message between agents

    agent-tester to agent-dispatcher

    via comment@mention

    “Nightly e2e on main (a41c9d7) failed in 2 of 48 cases: discount + VAT in the cart. The run report is attached.”

    Step 3 of 12 node: Dispatcher

    The dispatcher opens issue #251

    1. status: new
    2. status: in progress
    3. review
    4. testing
    5. awaiting approval
    What happens

    The dispatcher turns the report into a task labelled bug, backend, status: new and night, and assigns it to the backend developer. Nobody gets woken up.

    On GitHub

    Cart: VAT is off with a discount (rounding) #251

    Open agent-dispatcher opened this issue

    Labels
    bugbackendstatus: newnight
    Assignees
    agent-backend

    agent-dispatcher commented · 01:12

    Found by the nightly test (2 of 48). With a 10% discount and 21% VAT, the cart VAT differs from the invoice by 0.01 CZK. Rule: AGENTS.md → Prices. @agent-backend please take it.

    Message between agents

    agent-dispatcher to agent-backend

    via new issuelabelsassignment@mention

    “Found by the nightly test (2 of 48). With a 10% discount and 21% VAT, the cart VAT differs from the invoice by 0.01 CZK. Please take it.”

    Step 4 of 12 node: Task queue

    A developer picks the task from the queue

    1. status: new (done)
    2. status: in progress
    3. review
    4. testing
    5. awaiting approval
    What happens

    The task waits in the queue until an agent with the right domain takes it. Taking it means swapping a label and posting a short comment with a plan.

    On GitHub

    Cart: VAT is off with a discount (rounding) #251

    Open agent-dispatcher opened this issue

    Labels
    bugbackendstatus: in progressnight
    Assignees
    agent-backend

    agent-backend added status: in progress and removed status: new labels · 01:14

    agent-backend commented · 01:14

    On it. Plan: 1) in vat.ts, round the discounted base before calculating VAT, 2) tests for both cases from the night run. Branch fix/251-vat-discount.

    Message between agents

    agent-backend to agent-dispatcher

    via labelcomment

    “On it. Plan: 1) in vat.ts, round the discounted base before calculating VAT, 2) tests for both cases from the night run.”

    Step 5 of 12 node: Developers

    Pull request #252 and green CI

    1. status: new (done)
    2. status: in progress (done)
    3. review
    4. testing
    5. awaiting approval
    What happens

    The fix is built in its own branch; main does not change. Once CI passes, the developer requests a review and switches the label to review.

    On GitHub

    fix(cart): VAT from the rounded base after discount #252

    Open agent-backend wants to merge 1 commit: fix/251-vat-discount into main

    backendreviewnight

    Fixes #251

    src/pricing/vat.ts 3e8d1a4

    @@ -12,3 +12,4 @@export function vatFromDiscounted(base: number, discount: number, rate: number) {removed:   return round2(base * (1 - discount) * rate);added:   const discounted = round2(base * (1 - discount));added:   return round2(discounted * rate);}

    All checks have passed

    • passed: ci / build 52 s
    • passed: ci / unit 211 of 211
    • passed: ci / e2e cart 48 of 48

    agent-backend added review and removed status: in progress labels · 01:41
    requested a review from agent-reviewer

    Message between agents

    agent-backend to agent-reviewer

    via pull requestreview labelreview request

    “Fixes #251. CI is green, please review.”

    Step 6 of 12 node: Reviewer

    Review: changes requested

    1. status: new (done)
    2. status: in progress (done)
    3. review
    4. testing
    5. awaiting approval
    What happens

    The code is read by a different agent than the one who wrote it. The fix is correct, but a test for the rounding edge case is missing, so it goes back.

    On GitHub

    agent-reviewer requested changes · 01:52 Changes requested

    src/pricing/vat.ts

    @agent-backend the fix is correct, but there is no test for an amount that rounds exactly at the 0.005 CZK boundary. I won't pass it on without one.

    Message between agents

    agent-reviewer to agent-backend

    via review@mention

    “The fix is correct, but there is no test for an amount that rounds exactly at the 0.005 CZK boundary. I won't pass it on without one.”

    Step 7 of 12 node: Developers

    Fix and a new test

    1. status: new (done)
    2. status: in progress (done)
    3. review
    4. testing
    5. awaiting approval
    What happens

    The author addresses the comment right away, without waiting for the morning. The new commit reruns CI and previous approvals are dismissed.

    On GitHub

    agent-backend pushed 5be0c2f test(cart): 0.005 CZK rounding boundary · 02:08

    src/pricing/vat.test.ts 5be0c2f

    @@ -18,3 +18,8 @@test('VAT from the rounded base after a 10% discount', () => {  expect(vatFromDiscounted(1234.5, 0.1, 0.21)).toBe(233.32);});added: added: test('0.005 CZK rounding boundary', () => {added:   // base after discount 10.005 CZK → 10.01 CZK → VAT 2.10 CZKadded:   expect(vatFromDiscounted(20.01, 0.5, 0.21)).toBe(2.1);added: });

    All checks have passed

    • passed: ci / build 50 s
    • passed: ci / unit 212 of 212 (+1)
    • passed: ci / e2e cart 48 of 48

    agent-backend commented · 02:10

    @agent-reviewer added in 5be0c2f: a test for the 0.005 CZK boundary. Please review again.

    Message between agents

    agent-backend to agent-reviewer

    via comment@mention

    “Added in 5be0c2f: a test for the 0.005 CZK boundary. Please review again.”

    Step 8 of 12 node: Reviewer

    Review: approved

    1. status: new (done)
    2. status: in progress (done)
    3. review (done)
    4. testing
    5. awaiting approval
    What happens

    The second review passes. The reviewer swaps the review label for test, handing the work over to the tester.

    On GitHub

    agent-reviewer approved these changes · 02:15 Approved

    Code looks good, the AGENTS.md → Prices rule is followed. @agent-tester handing over for testing.

    agent-reviewer added testing and removed review labels · 02:15

    Message between agents

    agent-reviewer to agent-tester

    via reviewtest label@mention

    “Code looks good, the AGENTS.md → Prices rule is followed. Handing over for testing.”

    Step 9 of 12 node: Tester

    The tester verifies the fix

    1. status: new (done)
    2. status: in progress (done)
    3. review (done)
    4. testing
    5. awaiting approval
    What happens

    The tester runs exactly the two cases that failed overnight, plus the full cart suite, and compares screenshots too.

    On GitHub

    agent-tester · e2e on 5be0c2f

    • passed: cart: 48 of 48 tests
    • passed: 10% discount + 21% VAT
    • passed: 15% discount + 12% VAT
    • passed: cart summary screenshots no changes

    agent-tester commented · 02:40

    Verified on 5be0c2f: both night cases pass, full cart suite 48 of 48, screenshots unchanged. @agent-devops please deploy a preview.

    Message between agents

    agent-tester to agent-devops

    via comment@mention

    “Verified on 5be0c2f: both night cases pass, full cart suite 48 of 48, screenshots unchanged. Please deploy a preview.”

    Step 10 of 12 node: DevOps

    Preview and hand-over to a human

    1. status: new (done)
    2. status: in progress (done)
    3. review (done)
    4. testing (done)
    5. awaiting approval
    What happens

    DevOps deploys a preview with the fix; production stays unchanged. The dispatcher switches the label to awaiting approval and the night's work is done.

    On GitHub

    deploy-preview · Success
    pull_request · #252 · fix/251-vat-discount @ 5be0c2f

    1. passed: build 49 s
    2. passed: preview upload 8 s
    3. passed: preview smoke test 12 of 12

    agent-devops deployed to Preview · 02:45 · 5be0c2f

    Active View deployment

    agent-dispatcher added awaiting approval and removed testing labels · 02:46

    agent-dispatcher commented · 02:46

    @product-owner done and tested, the preview runs on 5be0c2f. Only approval and merge are left for the morning.

    Message between agents

    agent-devops to agent-dispatcher

    via deployment

    “Preview 5be0c2f deployed, smoke test 12 of 12.”

    agent-dispatcher to product-ownerhuman

    via labelcomment@mention

    “Done and tested, the preview runs on 5be0c2f. Only approval and merge are left for the morning.”

    Step 11 of 12 · morning node: Human

    Morning: a human approves and merges

    1. status: new (done)
    2. status: in progress (done)
    3. review (done)
    4. testing (done)
    5. awaiting approval
    What happens

    The product owner checks the preview and approves the PR. Merging to main is the one step agents are not allowed to do.

    On GitHub

    fix(cart): VAT from the rounded base after discount #252

    Merged agent-backend merged 2 commits: fix/251-vat-discount into main

    backendawaiting approvalnight

    product-owner approved these changes · 07:42 Approved

    Checked the preview, VAT matches in both the cart and the invoice. Approved.

    Merged product-owner merged commit 9c27f0b into main from fix/251-vat-discount · 07:43

    Message between agents

    product-ownerhuman to agent-devops

    via reviewmerge to main

    “Checked the preview, VAT matches in both the cart and the invoice. Approved.”

    Step 12 of 12 · morning node: Production

    Production deployed, issue closed

    1. status: new (done)
    2. status: in progress (done)
    3. review (done)
    4. testing (done)
    5. awaiting approval (done)
    6. closed ✓
    What happens

    The merge to main triggers the production deployment and issue #251 closes automatically. Every deployment can be traced to a commit and rolled back.

    On GitHub

    Cart: VAT is off with a discount (rounding) #251

    Closed agent-dispatcher opened this issue

    Labels
    bugbackendnight
    Assignees
    agent-backend

    product-owner closed this as completed in #252 · 07:43

    deploy-production · Success
    push · main · main @ 9c27f0b

    1. passed: build 51 s
    2. passed: deployment 14 s
    3. passed: production smoke test 12 of 12

    agent-devops deployed to Production · 07:47 · 9c27f0b

    Active View deployment

    agent-devops commented · 07:48

    Production runs on 9c27f0b, smoke test 12 of 12. The previous version a41c9d7 is ready for rollback.

    Message between agents

    agent-devops to agent-dispatcher

    via deploymentcomment

    “Production runs on 9c27f0b, smoke test 12 of 12. The previous version a41c9d7 is ready for rollback.”

    Roles in the diagram
    1. Human priorities, approval

      Sets priorities and requirements, approves merges to production and manages access, licenses and secrets. In disputed matters the human decides, not an agent.

    2. Dispatcher planning

      Turns requirements into concrete tasks, assigns labels and watches the dependencies between them. Handles tester reports and decides what gets fixed now and what can wait.

    3. Task queue GitHub Issues

      The only place where tasks exist. Each issue has a domain label, a status and a discussion history. Agents take work from here, not from chat.

    4. Developers multiple instances

      Several instances, each for its own domain (e.g. graphics, terrain, game logic, UI design). They work in their own branch and deliver the result as a pull request.

    5. Reviewer code review

      Reads the pull request as a different instance than the one that wrote the code. Checks compliance with AGENTS.md, bugs, security and needless complexity. Can send the change back.

    6. Tester Playwright

      Runs automated tests and walks through the application in a browser, compares screenshots. Files the bugs it finds as new issues and returns them to the dispatcher.

    7. DevOps CI, deployment

      Maintains the CI pipeline, builds and deployments. Watches logs and the state of environments. A change goes to production only after human approval.

    8. Production Cloudflare

      The deployed version after a merge to main. Every deployment can be traced to a specific commit and rolled back.

    This is how it looks in our repo

    Video: one day of our AI agent team From a task on GitHub through a pull request and automatic checks to a deployment on Cloudflare and an app from GitLab running in Azure. Real records, 70 s.

    Real screenshots from the repository of the game Rónin, developed by a team of AI agents. The agents work in Czech; GRAFIKA (graphics) and TEREN (terrain) are session names. Names of people, avatars and addresses are hidden.

    • The dispatcher issues a rule for all sessions as an issue with domain labels, not as a chat message.
      The dispatcher issues a rule for all sessions as an issue with domain labels, not as a chat message.
    • A task from testing: priority, owner (the GRAFIKA session) and a checklist of findings.
      A task from testing: priority, owner (the GRAFIKA session) and a checklist of findings.
    • The GRAFIKA session reports the status of its fixes with links to PRs and commits, and confirms after the merge.
      The GRAFIKA session reports the status of its fixes with links to PRs and commits, and confirms after the merge.
    • A pull request from the TEREN session: its own branch, changes described level by level, checks and the number of changed files.
      A pull request from the TEREN session: its own branch, changes described level by level, checks and the number of changed files.
    • The session reports “done in PR”, auto-merge merges the PR and the issue closes itself as completed.
      The session reports “done in PR”, auto-merge merges the PR and the issue closes itself as completed.

    Take it into your own repo

    The same rules and templates our agent team works by. The templates are in Czech. CI and settings files have a .txt extension so they do not run by accident. After downloading, rename them according to the “Where to” line.

    • AGENTS.md

      Shared rules for all agents: taking a task, branches, PRs, review, tests, night mode and what to do when they get stuck.

      Where to: AGENTS.md (repository root)

      Download
      File preview (51 lines)
      # AGENTS.md
      
      Pravidla pro všechny AI agenty v tomto repozitáři. Přečti je před každým úkolem.
      Agenti spolu mluví jen přes GitHub: štítky, komentáře, review a @zmínky. Chat se nepočítá.
      
      ## Převzetí úkolu
      - Ber jen issue se štítkem své domény (`backend`, `frontend`, `ui`, …) a `stav: nový`.
      - Při převzetí vyměň štítek `stav: nový` za `stav: v práci` a napiš komentář
        „Beru.“ s krátkým plánem (2–4 body).
      - Na jednom úkolu pracuje jeden agent. Má-li issue přiřazeného někoho jiného, neber ho.
      
      ## Větve a commity
      - Jedna větev na úkol: `fix/<číslo>-<popis>` nebo `feat/<číslo>-<popis>`,
        např. `fix/251-dph-sleva`.
      - Commity podle Conventional Commits: `fix(kosik): …`, `test(kosik): …`.
      - Nikdy nepiš přímo do `main` a nepoužívej force push na cizí větve.
      
      ## Pull request
      - Popis podle šablony PR, vždy s řádkem „Opravuje #číslo“.
      - O review požádej až při zelené CI: štítek `review` a žádost o review
        pro `agent-reviewer`.
      - Review dělá vždy jiný agent, než který kód napsal. Vlastní PR neschvaluj.
      - Po „Changes requested“ oprav, pushni a odpověz komentářem s @zmínkou reviewera.
      
      ## Testy
      - Každá oprava chyby má test, který bez opravy selže.
      - Po schválení reviewerem převezme PR `agent-tester` (štítek `test`).
      - Tester ověří přesně případ z úkolu a pak celou sadu dané oblasti.
      
      ## Merge a produkce
      - Do `main` nikdy nemerguj. Merge dělá jen člověk po vlastním schválení.
      - Hotový PR označ štítkem `čeká na schválení` a @zmiň vlastníka produktu.
      - `agent-devops` nasazuje jen náhledy. Produkce se nasazuje až po merge.
      
      ## Když si nevíš rady
      - Rozpor v zadání, obchodní rozhodnutí nebo chybějící přístup:
        štítek `blokováno`, komentář s konkrétní otázkou a @zmínka `agent-dispecer`.
      - Pak pokračuj jiným úkolem. Nečekej a nehádej.
      - Po dvou neúspěšných pokusech o stejnou opravu přestaň a nahlas `blokováno`.
      
      ## Noční režim (štítek `noc`)
      - V noci ber jen úkoly se štítkem `noc`, které večer schválil člověk,
        a chyby, které v noci najdou testy.
      - Žádné změny závislostí, databázových schémat ani infrastruktury.
      - Nejvýš 3 otevřené PR najednou a nejvýš 5 běhů CI na jeden PR.
      - Ráno musí mít každý úkol jasný stav: štítek a poslední komentář.
      
      ## Zakázáno
      - Měnit secrets, nastavení repozitáře, pravidla větví a CODEOWNERS.
      - Sahat na produkční data.
      - Vypínat nebo přeskakovat testy, aby CI prošla.
    • dispecer-prompt.md

      Instructions for the dispatcher: creating tasks, keeping work flowing, the night run and the morning summary for a human.

      Where to: dispatcher agent instructions

      Download
      File preview (43 lines)
      # Instrukce pro agenta-dispečera
      
      Jsi `agent-dispecer`. Neprogramuješ. Převádíš zadání a hlášení na úkoly,
      přiděluješ je a hlídáš, aby se práce nezasekla. Pracuješ jen přes GitHub.
      
      ## Vstupy
      - Nové issue a komentáře s @zmínkou `agent-dispecer`.
      - Hlášení od `agent-tester` (neúspěšné běhy testů).
      - Štítek `blokováno` u kteréhokoli úkolu.
      
      ## Založení úkolu
      1. Použij šablonu úkolu. Název: „Oblast: co je špatně“.
      2. Štítky: druh (`chyba` / `funkce`), doména, `stav: nový`.
      3. Přiřaď vývojáře podle domény. Nevíš-li, kterou doménu zvolit, zeptej se člověka.
      4. Do komentáře napiš, co přesně ověřit, a @zmiň přiřazeného agenta.
      5. Duplicitní hlášení nezakládej znovu, připoj ho komentářem k existujícímu úkolu.
      
      ## Hlídání toku
      - Úkol `stav: v práci` bez nového komentáře déle než 2 hodiny: zeptej se autora.
      - PR se štítkem `review` bez review déle než 1 hodinu: @zmiň `agent-reviewer`.
      - Úkol `blokováno`: jde-li rozhodnout podle AGENTS.md, rozhodni a napiš proč.
        Jinak @zmiň vlastníka produktu a nech úkol ležet.
      - Po zeleném testu a nasazeném náhledu vyměň štítek `test` za `čeká na schválení`
        a @zmiň vlastníka produktu jedním shrnujícím komentářem.
      
      ## Noční běh
      - Večer člověk schválí frontu štítkem `noc`. Pracuje se jen na těchto úkolech
        a na chybách, které v noci najdou testy.
      - Chyba z nočního testu: založ úkol se štítky `chyba`, doména, `stav: nový`, `noc`.
      - Nejvýš 3 rozpracované PR najednou. Další úkoly počkají ve frontě.
      - Nic nemerguj a nic nenasazuj do produkce. Náhledy nasazuje `agent-devops`.
      - Když se oprava dvakrát vrátí z review nebo z testu, označ ji `blokováno`
        a nech ji na ráno.
      
      ## Ranní shrnutí (do 7:00)
      Jeden komentář v úkolu noční fronty:
      - co je hotové a čeká na schválení (čísla PR),
      - co je blokované a proč (jedna věta, konkrétní otázka),
      - co zůstalo ve frontě.
      
      ## Nikdy
      - Neměň kód, testy, secrets ani nastavení repozitáře.
      - Neodpovídej za člověka a neschvaluj za něj.
    • ISSUE_TEMPLATE-ukol.md

      A task template that tells the agent what to fix and when it is done.

      Where to: .github/ISSUE_TEMPLATE/ukol.md

      Download
      File preview (36 lines)
      ---
      name: Úkol pro agenty
      about: Zadání, které si převezme agent podle štítku své domény
      title: "[oblast]: krátký popis"
      labels: ["stav: nový"]
      assignees: []
      ---
      
      <!--
        Šablona patří do .github/ISSUE_TEMPLATE/ukol.md.
        Štítek domény (backend, frontend, ui, …) doplní dispečer nebo člověk.
        Úkol pro noční běh dostane navíc štítek „noc“. Ten schvaluje člověk.
      -->
      
      ## Co se děje
      <!-- Jedna až dvě věty. U chyby: co je špatně a kde. -->
      
      ## Jak zopakovat
      1.
      2.
      3.
      
      ## Očekáváno
      <!-- Jak se má aplikace chovat. Odkaz na pravidlo v AGENTS.md, pokud existuje. -->
      
      ## Hotovo, když
      - [ ] test, který bez opravy selže
      - [ ] CI zelená (build, unit testy, e2e)
      - [ ] review od jiného agenta
      - [ ] náhled nasazený, štítek `čeká na schválení`
      
      ## Omezení
      <!-- Co se měnit nesmí (API, databáze, závislosti). Nepovinné. -->
      
      ## Příloha
      <!-- Výstup testu, log, screenshot. Žádné osobní ani produkční údaje. -->
    • PULL_REQUEST_TEMPLATE.md

      A pull request template linked to the task, with a checklist whose last item belongs to a human.

      Where to: .github/PULL_REQUEST_TEMPLATE.md

      Download
      File preview (27 lines)
      <!--
        Šablona patří do .github/PULL_REQUEST_TEMPLATE.md.
        Název PR podle Conventional Commits,
        např. „fix(kosik): DPH ze zaokrouhleného základu po slevě“.
      -->
      
      Opravuje #
      
      ## Co se mění
      <!-- Soubory a důvod, jeden bod na změnu. -->
      -
      
      ## Jak jsem to ověřil
      - [ ] nový test, který bez opravy selže
      - [ ] unit testy zelené
      - [ ] e2e testy dotčené oblasti zelené
      - [ ] náhled nasazený (odkaz doplní agent-devops v deploymentu)
      
      ## Rizika
      <!-- Co se může rozbít jinde. „Žádná známá“ je platná odpověď. -->
      
      ## Kontrolní seznam
      - [ ] větev `fix/<číslo>-<popis>` nebo `feat/<číslo>-<popis>`
      - [ ] žádné změny secrets, závislostí ani nastavení repozitáře
      - [ ] review od jiného agenta, než je autor
      - [ ] štítek `čeká na schválení` až po zeleném testu a nasazeném náhledu
      - [ ] schválení a merge člověkem (agenti tento bod neodškrtávají)
    • stitky.yml.txt

      Labels the agents use to hand work over: statuses, work types, domains and night.

      Where to: .github/labels.yml

      Download
      File preview (50 lines)
      # Definice štítků pro repozitář s týmem AI agentů.
      # Formát odpovídá běžným nástrojům pro synchronizaci štítků.
      # Soubor ulož jako .github/labels.yml (bez přípony .txt).
      
      # --- Stav úkolu: agenti si předávají práci změnou štítku ---
      - name: "stav: nový"
        color: "ededed"
        description: "Úkol čeká, až si ho převezme agent své domény"
      - name: "stav: v práci"
        color: "fbca04"
        description: "Agent úkol převzal a napsal plán do komentáře"
      - name: "review"
        color: "1d76db"
        description: "PR čeká na review od jiného agenta"
      - name: "test"
        color: "0e8a16"
        description: "Review prošlo, na řadě je agent-tester"
      - name: "čeká na schválení"
        color: "8a5200"
        description: "Otestováno, náhled běží, rozhoduje člověk"
      - name: "blokováno"
        color: "b60205"
        description: "Agent potřebuje rozhodnutí nebo přístup, viz poslední komentář"
      
      # --- Druh práce ---
      - name: "chyba"
        color: "d73a4a"
        description: "Něco nefunguje podle zadání"
      - name: "funkce"
        color: "a2eeef"
        description: "Nová funkce nebo rozšíření"
      
      # --- Domény: určují, který vývojář úkol převezme ---
      - name: "backend"
        color: "5319e7"
        description: "Server, API, výpočty cen"
      - name: "frontend"
        color: "0052cc"
        description: "Webové rozhraní a klientský kód"
      - name: "ui"
        color: "d876e3"
        description: "Návrh rozhraní, texty, přístupnost"
      - name: "infra"
        color: "475569"
        description: "CI, nasazení, prostředí (jen agent-devops)"
      
      # --- Plánování ---
      - name: "noc"
        color: "1f2a44"
        description: "Schváleno člověkem pro noční běh agentů"
    • ci.yml.txt

      GitHub Actions for every pull request: build, unit tests and Playwright e2e tests.

      Where to: .github/workflows/ci.yml

      Download
      File preview (60 lines)
      # Vzor GitHub Actions: kontroly u každého pull requestu.
      # Soubor ulož jako .github/workflows/ci.yml (bez přípony .txt).
      # Názvy úloh musí sedět s povinnými kontrolami v pravidlech větve main.
      name: ci
      
      on:
        pull_request:
          branches: [main]
      
      permissions:
        contents: read
      
      concurrency:
        group: ci-${{ github.head_ref }}
        cancel-in-progress: true   # nový push zruší starý běh a šetří minuty
      
      jobs:
        build:
          runs-on: ubuntu-latest
          steps:
            - uses: actions/checkout@v4
            - uses: actions/setup-node@v4
              with:
                node-version: 22
                cache: npm
            - run: npm ci
            - run: npm run build
      
        unit:
          runs-on: ubuntu-latest
          steps:
            - uses: actions/checkout@v4
            - uses: actions/setup-node@v4
              with:
                node-version: 22
                cache: npm
            - run: npm ci
            - run: npm test                   # unit testy
      
        e2e:
          needs: build
          runs-on: ubuntu-latest
          timeout-minutes: 20
          steps:
            - uses: actions/checkout@v4
            - uses: actions/setup-node@v4
              with:
                node-version: 22
                cache: npm
            - run: npm ci
            - run: npx playwright install --with-deps chromium
            # Server pro testy spouští playwright.config (sekce webServer).
            - run: npx playwright test
            # Při selhání uložíme report, aby ho agent-tester mohl přiložit k úkolu.
            - if: failure()
              uses: actions/upload-artifact@v4
              with:
                name: playwright-report
                path: playwright-report/
                retention-days: 7
    • ochrana-main.yml.txt

      Rules for the main branch: one required human approval and required CI checks.

      Where to: Settings → Rules (set up manually)

      Download
      File preview (36 lines)
      # Popis pravidel pro větev main (GitHub → Settings → Rules → Rulesets).
      # GitHub tento soubor nenačte sám: nastavení zadej ručně nebo přes API.
      # Měnit ho smí jen člověk s rolí správce, nikdy agent.
      ruleset:
        name: ochrana-main
        target: branch
        include: ["refs/heads/main"]
        enforcement: active
      
        rules:
          # Nic se nezapisuje přímo do main, jen přes pull request.
          pull_request:
            required_approving_review_count: 1     # 1 schválení od člověka
            require_code_owner_review: true        # vlastníci kódu = jen lidé (CODEOWNERS)
            dismiss_stale_reviews_on_push: true    # nový commit = nové schválení
            require_last_push_approval: true       # poslední push schvaluje někdo jiný
            required_review_thread_resolution: true
      
          # Povinné kontroly: názvy úloh z .github/workflows/ci.yml
          required_status_checks:
            strict: true                           # větev musí být aktuální vůči main
            checks:
              - ci / build
              - ci / unit
              - ci / e2e
      
          non_fast_forward: true                   # zákaz force push
          deletion: true                           # main nejde smazat
          required_linear_history: true            # merge jen jako squash
      
        # Výjimky nemá nikdo: ani správci, ani účty agentů.
        bypass_actors: []
      
      # Účty agentů (GitHub App nebo strojové účty) mají v repozitáři roli „Write“.
      # Schválení od agent-reviewer je kontrola kvality, ale nesplní code owner review.
      # Merge do main proto vždy dělá člověk.
    • CODEOWNERS.txt

      Only people are code owners, so an approval from an agent is not enough to merge.

      Where to: .github/CODEOWNERS

      Download
      File preview (20 lines)
      # Vlastníci kódu. Soubor ulož jako .github/CODEOWNERS (bez přípony .txt).
      # Uvádějte jen lidi nebo týmy lidí, nikdy účty agentů.
      # Díky tomu pravidlo „require code owner review“ vždy vyžaduje schválení člověkem.
      #
      # Formát: <vzor cesty> <tým nebo účet>
      # Platí poslední odpovídající řádek.
      
      # Výchozí vlastník celého repozitáře
      *                         @organizace/vlastnici-produktu
      
      # Ceny, DPH a platby: schvalují lidé, kteří rozumí účetním pravidlům
      /src/pricing/             @organizace/vlastnici-produktu @organizace/ucetni-pravidla
      /src/payments/            @organizace/vlastnici-produktu @organizace/ucetni-pravidla
      
      # Pravidla pro agenty a nastavení repozitáře mění jen správci
      /AGENTS.md                @organizace/spravci
      /.github/                 @organizace/spravci
      
      # Infrastruktura a nasazení
      /infra/                   @organizace/spravci

    Agent communication

    Communication

    How the agents talk to each other

    Four model situations step by step: who hands what to whom, and where a human steps in. Nothing goes to production without their approval.

    Rule: a change goes to production only after a human approves it. Agents prepare, review and test.

    Brief:

      Bug fix: “The form can't be submitted on mobile.”

      1. 09:02 · Human → Dispatcher · Comment The form can't be submitted on mobile. Customers are complaining. A person describes the problem in their own words. Nothing more is needed.
      2. 09:04 · Dispatcher → Developer (frontend) · Issue Creating the issue “Form can't be submitted on mobile”. The frontend developer takes it, priority is high. The Dispatcher turns the report into a specific issue and decides who takes it and how urgent it is.
      3. 09:31 · Developer (frontend) → Reviewer · Pull request Done. On small screens the button was covered by another element. Sending it for review. The Developer works on its own copy of the code. It submits the change for review and cannot push it to production itself.
      4. 09:40 · Reviewer → Developer (frontend) · Review The fix works, but there's no test that would catch the same bug next time. Sending it back. The review is done by a different agent than the author. It sent the change back because nothing prevented the bug from coming back.
      5. 09:52 · Developer (frontend) → Reviewer · Pull request The test is added and the automated checks passed. The author addresses the comment and resubmits the change.
      6. 09:55 · Reviewer → Tester · Review Now it's fine. Approved from a code perspective, passing it to the Tester. The code review passed. But the change is not in production yet.
      7. 10:10 · Tester → DevOps · Test Tested in the browser on three mobile sizes. The form submits. The Tester tries the fix the way a customer would use it.
      8. 10:14 · DevOps → Human · Deployment The fix is running on a staging URL and awaits your approval. The staging version (preview) is separate from production; customers don't see it.
      9. 13:05 · Human → DevOps · Approval I've checked the preview, approved. This is the only step that releases the change to production. It is always done by a human.
      10. 13:07 · DevOps → Production · Deployment Deployed to production; the post-deployment check is fine. Deployment runs automatically only after approval and can be rolled back to the previous version.
      • Agents involved: 5 (Dispatcher, Developer (frontend), Reviewer, Tester, DevOps)
      • Human stepped in 2×: brief, approval
      • Checks before production deployment: 5 (2× automated tests, 2× code review, 1× browser test)

      New feature: “Export invoices to Excel.”

      1. 08:30 · Human → Dispatcher · Comment We need to export invoices to Excel for a selected period and customer. The brief can be written as an ordinary request.
      2. 08:34 · Dispatcher → Architect · Issue I've split the brief into three issues: interface spec, file generation, and a button with a filter. Architect, please design the interface. A large brief is split into smaller issues that can run in parallel.
      3. 08:50 · Architect → Developers · Comment The export takes a period and a customer and returns an Excel file. You both build on this spec. The Architect decides up front how the parts fit together, so the two developers don't work against each other.
      4. 10:20 · Developer (backend) → Reviewer · Pull request File generation is done. Backend and frontend are built in parallel, each in its own branch.
      5. 10:45 · Developer (frontend) → Reviewer · Pull request The button and filter are done. The second part of the work finished independently of the first.
      6. 11:05 · Reviewer → Tester · Review Both changes match the interface spec. Approved from a code perspective. The Reviewer also checks that both parts follow the shared spec.
      7. 11:40 · Tester → Dispatcher · Test An export with more than 1,000 invoices fails. Returning it to the Dispatcher. The Tester found the bug before a customer did. It reports it to the Dispatcher, who decides the priority.
      8. 11:44 · Dispatcher → Developer (backend) · Issue New issue: large exports fail. It blocks the release, handle it first. The bug comes back as a new issue with a priority. Nothing gets lost in a chat.
      9. 13:10 · Developer (backend) → Reviewer · Pull request The export is now processed in batches, so large volumes go through. The fix goes through all the checks again, no shortcuts.
      10. 13:25 · Reviewer → Tester · Review The fix is fine. Approved from a code perspective. Even a small fix is reviewed by another agent.
      11. 13:50 · Tester → DevOps · Test The export of 5,000 invoices went through and the file content is correct. The Tester verifies exactly the case that failed before.
      12. 13:55 · DevOps → Human · Deployment The export is on the staging URL. It awaits your approval. A person can try the feature before customers see it.
      13. 15:00 · Human → DevOps · Approval Tested, approved. Without this approval the feature would stay on the staging URL only.
      14. 15:02 · DevOps → Production · Deployment Deployed to production. The new feature is available to customers. The previous version stays ready for a rollback.
      • Agents involved: 7 (Dispatcher, Architect, Developer (backend), Reviewer, Developer (frontend), Tester, DevOps)
      • Human stepped in 2×: brief, approval
      • Checks before production deployment: 7 (3× automated tests, 2× code review, 2× browser test)

      Human decision: “Two agents run into a conflict in the brief.”

      1. 10:02 · Developer (frontend) → Dispatcher · Comment Issue #240 says to show prices incl. VAT, issue #241 says excl. VAT for business customers. Both can't be met on one page. The agent notices that two briefs contradict each other and doesn't resolve it on its own.
      2. 10:04 · UI Designer → Dispatcher · Comment Confirmed. The screen design has room for one price and I need to know which one to show. A second agent hit the same problem from its side.
      3. 10:06 · Dispatcher → Human · Label Conflict in the brief. Agents shouldn't decide about prices, handing it to a human. It's a business decision, so the agents don't guess. The issue is flagged and waits for a human.
      4. 10:07 · Dispatcher → Developer (frontend) · Issue Until the decision arrives, continue with the issue “Faster loading of the order list”. Work doesn't stop; the agents take on other issues in the meantime.
      5. 11:30 · Human → Dispatcher · Comment Businesses excl. VAT, everyone else incl. VAT. Decide by account type. One sentence is enough. The decision is recorded on the issue and applies to further work too.
      6. 11:32 · Dispatcher → Developer (frontend) · Label Decided. The brief is updated, carry on. The issue goes back to the queue with a clear brief.
      7. 12:40 · Developer (frontend) → Reviewer · Pull request Prices are shown according to the account type. The Developer builds exactly on the decision the human recorded.
      8. 12:55 · Reviewer → Tester · Review The code matches the decision on the issue. Approved from a code perspective. The Reviewer compares the change with what the human decided.
      9. 13:20 · Tester → DevOps · Test Verified for both business and standard accounts. The Tester tries both variants the human decided on.
      10. 13:25 · DevOps → Human · Deployment The change is on the staging URL. It awaits your approval. Even after the human's decision the same rule applies: preview first, then approval.
      11. 14:00 · Human → DevOps · Approval Approved. Nothing goes to production without a human's approval.
      12. 14:02 · DevOps → Production · Deployment Deployed to production. Customers see prices according to their account type.
      • Agents involved: 6 (Developer (frontend), Dispatcher, UI Designer, Reviewer, Tester, DevOps)
      • Human stepped in 2×: decision, approval
      • Checks before production deployment: 3 (1× automated tests, 1× code review, 1× browser test)

      Night run: “The agents work without a human; in the morning only the approval is waiting.”

      1. 01:10 · Tester → Dispatcher · Test The nightly order test failed: with a discount, the VAT calculation in the cart is off. Tests run on their own every night. An agent finds the bug, not a customer.
      2. 01:12 · Dispatcher → Developer (backend) · Issue Creating an issue to fix the VAT calculation with discounts and assigning it to the backend developer. The Dispatcher sets the priority and assigns the work on its own. Nobody gets woken up at night.
      3. 01:40 · Developer (backend) → Reviewer · Pull request Fix done: VAT is now calculated from the rounded base after the discount. The fix is made in a separate branch; the main version stays unchanged.
      4. 01:52 · Reviewer → Developer (backend) · Review Missing a test for edge-case rounding (0.005 CZK). Sending it back. The review is done by another agent. It sends the change back even though nobody is watching.
      5. 02:10 · Developer (backend) → Reviewer · Pull request Test added, all checks passed. The author addresses the comment on its own, without waiting for the morning.
      6. 02:15 · Reviewer → Tester · Review Approved from a code perspective. The second review passed.
      7. 02:40 · Tester → DevOps · Test The full suite of 48 order tests passed, including both discount cases. The Tester verifies exactly the case that failed overnight, plus everything else.
      8. 02:45 · DevOps → Dispatcher · Deployment A preview with the fix is running on the staging URL. The staging version is ready for the morning check. Production stays unchanged.
      9. 02:46 · Dispatcher → Human · Label Done and tested. In the morning it awaits approval to merge to production. This is where the night's work ends. Only a human releases the change to production.
      • Agents involved: 5 (Tester, Dispatcher, Developer (backend), Reviewer, DevOps)
      • Human stepped in 0×: the whole run went without them; the merge to production awaits their approval
      • Checks before production deployment: 6 (2× browser test, 2× automated tests, 2× code review)

      Times show an illustrative run; the scenarios are models.

      Glossary 8 terms without jargon
      Issue
      A record of one piece of work, with the brief, discussion and result in one place.
      Label
      A tag on an issue that shows who will take it and how urgent it is.
      Branch
      A separate copy of the code to work in without changing the working version.
      Pull request (PR)
      A request to accept changes from a branch into the main version, with an overview of what changes.
      Review
      A check of the changes by another agent or a human before they are accepted.
      CI
      Automated builds and tests that run on every change without human involvement.
      Deployment
      Publishing a new version: first on a staging URL, then for customers after approval.
      Merge
      Merging an approved change into the main version. Here it is always confirmed by a human.

      On GitHub

      For engineers

      The night run on GitHub: no human involved until approval

      The "Night run" scenario in concrete form: issue, pull request, review, CI and repository rules. Agents have their own machine accounts. Between 1:10 and 2:46 no human steps in, but the merge to production stays locked for a human.

      1. status: new
      2. status: in progress
      3. review
      4. CI green
      5. testing
      6. preview
      7. awaiting approval
      8. merge: human

      Issue #251

      The dispatcher opened the issue on its own from a nightly test failure. Labels set the domain, status and priority.

      #251  Cart: VAT is off with a discount (rounding)
      Labels: bug · backend · status: new · night
      Opened by: agent-dispatcher · Assigned to: agent-backend
      
      ## What happens
      The nightly e2e test failed in 2 of 48 cases. With a 10% discount and 21% VAT,
      the VAT in the cart summary differs by 0.01 CZK from the VAT on the invoice.
      
      ## How to reproduce
      1. Cart with an item of 1,234.50 CZK
      2. 10% discount
      3. Compare the VAT in the cart summary and on the invoice
      
      ## Expected
      VAT is calculated from the rounded base after discount (AGENTS.md → Prices).
      
      ## Done when
      - [ ] tests for both cases from the nightly run
      - [ ] cart e2e suite green
      - [ ] preview deployed, label "awaiting approval"

      Pull request #252

      The Developer opens a PR from its own branch. Agents cannot tick the last item (human approval).

      #252  fix(cart): VAT from the rounded base after discount
      fix/251-vat-discount → main · author: agent-backend · review: agent-reviewer
      
      Fixes #251
      
      ## Changes
      - `src/pricing/vat.ts` – the base is rounded after the discount, before calculating VAT
      - `src/pricing/vat.test.ts` – 3 new tests, including the 0.005 CZK boundary
      
      ## Verification
      - [x] unit tests 212/212
      - [x] e2e cart 48/48
      - [x] preview: fix-251-vat-discount
      - [ ] human approval (merge into main)

      Review

      The Reviewer is a different agent with its own account. It sends the change back with a concrete suggestion.

      # agent-reviewer · Changes requested · src/pricing/vat.ts
      # "Round the base after the discount separately and add a test for the 0.005 CZK boundary."
      
      export function vatFromDiscounted(base: number, discount: number, rate: number) {
        return round2(base * (1 - discount) * rate);
        const discounted = round2(base * (1 - discount));
        return round2(discounted * rate);
      }
      
      test('VAT from the rounded base after a 10% discount', () => {
        expect(vatFromDiscounted(1234.5, 0.1, 0.21)).toBe(233.32);
      });

      CI

      Checks on every PR and a nightly test run. A failure is automatically filed as a new issue.

      # .github/workflows/ci.yml – checks on every PR and a nightly e2e run
      name: ci
      on:
        pull_request:
        schedule:
          - cron: '0 1 * * *'   # every night at 1:00 UTC
      
      jobs:
        test:
          runs-on: ubuntu-latest
          steps:
            - uses: actions/checkout@v4
            - uses: actions/setup-node@v4
              with: { node-version: 22, cache: npm }
            - run: npm ci
            - run: npm test                       # unit tests
            - run: npx playwright install --with-deps chromium
            - run: npx playwright test            # e2e in the browser
            # A nightly failure automatically becomes a task for the dispatcher
            - if: failure() && github.event_name == 'schedule'
              run: gh issue create --title "Nightly e2e failed" --label "bug,night" --body-file e2e-report.md
              env:
                GH_TOKEN: ${{ github.token }}

      Protecting main

      Repository rules enforce human approval. Agents can neither bypass nor change them.

      # Rules for the main branch (GitHub → Settings → Rules)
      main:
        require_pull_request: true          # nothing is written directly to main
        required_approvals: 1
        require_code_owner_review: true     # approval must come from a code owner = a human
        dismiss_stale_approvals: true       # new commit = new approval
        required_status_checks:
          - ci / test                       # unit + e2e must be green
        block_force_push: true
        bypass: none                        # nobody has an exception, not even agents
      
      # .github/CODEOWNERS
      # * @company/management               # code owners are humans only
      #
      # Agents have their own machine accounts (GitHub App). Approval from the reviewer agent
      # is a quality check, but it does not satisfy "code owner review". So a human always merges.

      AGENTS.md

      Shared rules that every agent reads before starting work.

      # AGENTS.md (excerpt)
      
      ## Tasks
      - Only take tasks labelled with your domain and `status: new`.
      - When you take a task, change the label to `status: in progress` and post your plan as a comment.
      - One branch per task: `fix/<number>-<description>` or `feat/<number>-<description>`.
      
      ## Pull request
      - Description follows the template, always "Fixes #number".
      - Request a review only when CI is green.
      - Never merge into main. A human does that.
      
      ## Prices
      - VAT is calculated from the rounded base after discount, rounded to 2 decimal places.
      
      ## When you are stuck
      - Conflicting requirements or a business decision: label `waiting for human`
        and continue with another task.
      
      ## Forbidden
      - Changing secrets, repository settings, production data and branch rules.

      gh commands

      This is how agents work with GitHub from the terminal. Only a human runs the last command.

      # Dispatcher: turns a nightly failure into a task
      gh issue create --title "Cart: VAT is off with a discount (rounding)" \
        --label "bug,backend,status: new,night" --assignee agent-backend --body-file report.md
      
      # Developer: takes the task, fixes it and opens a PR
      gh issue edit 251 --remove-label "status: new" --add-label "status: in progress"
      git switch -c fix/251-vat-discount
      git commit -am "fix(cart): VAT from the rounded base after discount"
      git push -u origin fix/251-vat-discount
      gh pr create --base main --title "fix(cart): VAT from the rounded base after discount" --body "Fixes #251"
      
      # Reviewer: goes through the changes and decides
      gh pr diff 252
      gh pr review 252 --request-changes --body "Add a test for the 0.005 CZK boundary."
      gh pr review 252 --approve
      
      # Tester and DevOps: wait for checks, deploy a preview, label the PR
      gh pr checks 252 --watch
      gh pr edit 252 --add-label "awaiting approval"
      
      # The human in the morning – the one step agents are not allowed to take
      gh pr merge 252 --squash --delete-branch

      The same process on GitLab

      We develop the Doklady AI application on GitLab. Issues, merge requests, pipelines and human approval work the same way as on GitHub. The code below is a real excerpt from the application; the workflow around it is a model example.

      Merged

      QR payment only on Czech documents !214

      agent-backend wants to merge fix/412-qr-cs-only into main · Closes #412

      bugbackendawaiting approval

      Pipeline #8812 passed
      1. build backend:bootJarfrontend:vite build
      2. test backend:unit (+2)frontend:vitest
      3. e2e playwright: invoice with QR
      4. review preview environment

      IssuedInvoiceService.kt +5 −0

      --- a/backend/src/main/kotlin/…/service/IssuedInvoiceService.kt
      +++ b/backend/src/main/kotlin/…/service/IssuedInvoiceService.kt
      @@ -4848,11 +4848,17 @@
          /** [amount] defaults to the face value; the public pay page passes the remaining amount. */
          private fun IssuedInvoiceEntity.qrPaymentPayload(amount: BigDecimal = totalWithVat): String? {
              if (!requiresTransferAccount(paymentMethod)) return null
              // The QR code uses the SPAYD format (`SPD*1.0`) with a Czech IBAN — a CZECH standard.
              // Polish or German banking apps cannot read it, so on a foreign-language
              // document it would be a square that does nothing. The document language is a more
              // reliable hint than the currency: a Czech customer may also pay an invoice in euros.
              if (!language.trim().equals("CS", ignoreCase = true) && language.isNotBlank()) return null
              val iban = czechIban(bankAccountNumber.orEmpty(), bankCode.orEmpty()) ?: return null
              if (amount <= BigDecimal.ZERO) return null
      
              val vs = variableSymbol
                  ?.filter(Char::isDigit)
                  ?.ifBlank { null }
                  ?: CisloFaktury.variabilniSymbol(invoiceNumber).ifBlank { null }
      
              val parts = mutableListOf(
                  "SPD*1.0",
                  "ACC:$iban",
                  "AM:${amount.setScale(2, RoundingMode.HALF_UP).toPlainString()}",
                  "CC:${currency.uppercase()}"
              )
      1. agent-dispatcher

        Opened issue #412: the QR payment code is also printed on Polish invoices, and banks there cannot read it. Labels: bug, backend, status: new.

      2. agent-backend

        Taking #412 (status: in progress). Fix in !214: QR only on Czech documents; the document language decides, not the currency.

      3. agent-reviewer

        Missing a test for an EUR invoice in Czech, which should have the QR code. Sending it back.

      4. agent-backend

        Added 2 tests (CS + EUR → QR yes, PL → QR no). Pipeline green.

      5. agent-reviewer

        Approved. Label: testing.

      6. agent-tester

        E2E: PLN invoice without QR, CZK and EUR with QR. Label awaiting approval.

      7. product owner

        Approved, merging.

      Agent roles

      Agent roles

      Who does what in the team

      Each agent has a narrowly defined role and a matching type of model. An expensive model where decisions are made or code is written, a cheaper one for routine work.

      • Dispatcher

        • Breaks the brief down into tasks and assigns labels.
        • Tracks the queue, dependencies and blocked tasks.
        • Triages test reports by severity.
        Supports
        Coordination (shared cost)
        Model
        Mid-tier model. Reads a lot, writes little, low token volume.
        Approximate cost
        40 USD / mo · approx. 854 CZK
      • Architect

        • Designs the module structure and interfaces between domains.
        • Maintains AGENTS.md and technical decisions.
        • Assesses larger changes before anyone starts writing them.
        Supports
        Programmer
        Model
        The strongest available model with a long context. Called less often.
        Approximate cost
        100 USD / mo · approx. 2,136 CZK
      • Developer

        3 instances
        • Implements tasks in its domain on its own branch.
        • Writes tests for its changes and submits a pull request.
        • The number of instances matches the number of project domains.
        Supports
        Programmer
        Model
        A strong coding model. Accounts for most of the token usage.
        Approximate cost
        3 × 350 USD = 1,050 USD / mo · approx. 22,427 CZK
      • Reviewer

        • Reviews pull requests as a different instance than the one that wrote the code.
        • Looks for bugs, security issues and convention violations.
        • Sends changes back with concrete comments.
        Supports
        Programmer
        Model
        A strong model. Reads a lot of code, short output.
        Approximate cost
        120 USD / mo · approx. 2,563 CZK
      • UI Designer

        • Designs screens and components directly in code following the design system.
        • Reads Figma files and screenshots, prepares variants to choose from.
        • Keeps colors, typography and accessibility consistent.
        Supports
        UI designer (Figma)
        Model
        A strong multimodal model that reads both images and code.
        Approximate cost
        150 USD / mo · approx. 3,204 CZK
      • Tester

        • Runs tests and walks through the application in a browser (Playwright).
        • Compares screenshots and looks for regressions.
        • Files bugs it finds as issues for the dispatcher.
        Supports
        Tester
        Model
        Mid-tier model with tools. Cost grows with the number of test runs.
        Approximate cost
        180 USD / mo · approx. 3,845 CZK
      • DevOps

        • Maintains the CI pipeline, builds and deployment.
        • Watches logs, build failures and the state of environments.
        • Prepares a rollback when a deployment fails.
        Supports
        DevOps specialist
        Model
        A smaller model. The work is repetitive and well documented.
        Approximate cost
        70 USD / mo · approx. 1,495 CZK
      Total for 9 agents 1,710 USD / mo ≈ 36,524 CZK / mo
      The same number of people (9) 1,053,000 CZK / mo 117,000 CZK per person per month incl. employer contributions

      People cost roughly 28× more

      Costs are the default values from the calculator below. Converted at the Czech National Bank rate as of September 25, 2026: 1 USD = 21.359 CZK.

      The human role

      What stays with the human

      Agents do most of the routine work. They do not take over responsibility for the result, the money or access.

      • Priorities and requirements

        Decides what gets done and in what order. Writes requirements so that it can be verified when they are done.

      • Approving merges to production

        No change goes into main without human approval. An agent can prepare, test and recommend, but not deploy.

      • Licenses

        Reviews the licenses of libraries and resources the agents want to use and is accountable for them.

      • Secrets and access

        Manages keys, tokens and permissions. Agents only get the access their role requires.

      • Disputed matters

        When agents disagree or a task has no clear solution, a human decides and the decision is recorded on the task.

      Day and night mode

      Comparison of day and night mode
      Day Night
      Who is in charge A human; the dispatcher prepares the groundwork The dispatcher, following approved priorities
      What happens Reviewing results, approvals, new requirements Implementation, tests, fixing bugs found by tests
      Merge into main Yes, after approval No, pull requests wait for the morning
      When a problem occurs A human decides right away The task is marked as blocked and the agent moves on to another one

      Servers and deployment

      Deployment on servers

      The cell: a dispatcher and its team

      The basic unit is a cell: one dispatcher and a group of agents sharing a task queue and a repository. A bigger project does not get a bigger dispatcher, it gets another cell.

      Capacity of one server

      A server with 8 GB RAM handles 4–6 agents, one with 16 GB roughly 8–10. Agents spend most of their time waiting for the model to respond. Memory is used mainly by builds and browser tests.

      The CPU is less of a constraint than memory, so RAM size is what matters most.

      Bottlenecks

      • API rate limits. More agents on one account means more waiting, not more work.
      • Code conflicts. Agents working in the same part of the code get in each other’s way. Splitting work by domain helps.
      • Human capacity. Someone approves every merge to production. That is usually the first limit.

      Which servers the agents run on

      The AI model itself runs at the provider (Anthropic) and is included in the agent price. Your own server only runs what the agent needs to work: a copy of the repository, builds, tests and a browser for the Tester. That is why a cheap Linux server is enough. In one cell the server accounts for about 4% of the monthly cost of agents and servers; the rest is the cost of the models.

      Option Server Specs Agents Price / mo When to choose
      Smaller team Hetzner Cloud CPX32 4 vCPU (dedicated), 8 GB RAM, 160 GB disk 4–6 35.49 EUR ≈ 864 CZK A smaller cell or a pilot. A few parallel runs are enough for browser tests.
      Recommended default Hetzner Cloud CPX42 8 vCPU (dedicated), 16 GB RAM, 320 GB disk 8–10 69.49 EUR ≈ 1,692 CZK A full cell. Matches the default server price in the calculator.
      Budget Hetzner Cloud CX43 8 vCPU (shared), 16 GB RAM, 160 GB disk 8–10 15.99 EUR ≈ 389 CZK A cheaper option for a full cell. Performance is shared and may be slower at peak times.
      Enterprise Microsoft Azure B4ms 4 vCPU, 16 GB RAM, disk billed separately 8–10 140 USD ≈ 2,990 CZK The company already uses Azure, has a Microsoft agreement and identity management. Cheaper with a 1–3 year reservation.

      Prices excl. VAT, as of September 2026. Converted at the Czech National Bank rate as of September 25, 2026 (1 EUR = 24.350 CZK, 1 USD = 21.359 CZK). A cell with 8–10 agents needs a server with 16 GB RAM; 8 GB is enough for a smaller cell.

      What runs on the server

      • Ubuntu Server 24.04 LTS. Server operating system, no license fees.
      • Claude Code. The agent in the terminal. Its cost is the agent price in the calculator; the server does not add to it.
      • Git, GitHub CLI, Node.js. Repository work, builds and running tests.
      • Playwright with Chromium. The Tester uses it to walk through the application. It uses the most memory.
      • Docker. Each agent runs separately in its own container with its own copy of the repository.
      • Firewall and SSH keys. The server accepts no incoming connections except for administration. Production passwords and keys are never stored on it.

      Other running costs

      • Server backups (Hetzner): 20% of the server price. Optional; the code is stored on GitHub anyway.
      • GitHub: 0 CZK to start. Free private repositories and 2,000 CI minutes per month; paid plans from 4 USD per user.
      • Cloudflare Pages: 0 CZK. Hosting for websites and previews on the Free plan.
      • AI model (Claude): agent price. The largest item. The model runs at Anthropic; nothing is computed on the server.

      Deployment sizes

      Option Agents Composition Cells Servers People managing Agents + servers / mo Action
      Small 4 Dispatcher, 2× Developer, Tester Review, design and deployment stay with people or run automatically. A server with 8 GB RAM is enough. 1 1× 8 GB 1 ≈ 21,000 CZK
      Medium 8 Dispatcher, Architect, 2× Developer, Reviewer, UI Designer, Tester, DevOps A complete cell: every role is filled by one agent, plus a second developer. 1 1× 16 GB 1–2 ≈ 31,000 CZK
      Large 15 2× Dispatcher, Architect, 6× Developer, 2× Reviewer, UI Designer, 2× Tester, DevOps Two cells sharing an architect, a designer and DevOps. Each cell has its own dispatcher. 2 2× 16 GB 2 ≈ 70,000 CZK

      The cost does not include the people managing the team. The full calculation including them is in the calculator.

      Rónin case study

      Case study

      How we do it today: the Rónin project

      Rónin is a 3D game for web browsers, Android and iOS. It is developed by a group of specialized AI sessions coordinated by a dispatcher. A human sets the direction and approves what goes into production.

      Numbers and estimates

      Real data from the repository as of September 28, 2026. Costs and dates are our estimate.

      Running since
      September 23, 2026, i.e. 5.5 days non-stop including nights
      Agents (sessions)
      7: a dispatcher and six sessions by area
      Done
      1,044 commits on main, 254 merged pull requests, 82 closed tasks
      Platforms
      web browser, Android and iOS (mouse, keyboard and touch controls)
      AI agent teamThe same number of people (7)
      Monthly cost ≈ 46,000 CZK 819,000 CZK
      Cost of the work so far ≈ 8,000 CZK for 5.5 days ≈ 3,300,000 CZK for 4 months of work
      Playable version of the whole game target October 3–5, 2026 estimate 6–8 months
      First public demo for ordinary players estimate October 10–20, 2026 estimate 8–10 months
      Commercial product (web, Android, iOS) estimate January–February 2027, ≈ 210,000 CZK for the agents estimate 12–18 months, ≈ 10,000,000 CZK–15,000,000 CZK

      Agent cost is based on the role prices in the calculator (six sessions at the developer price plus the dispatcher), excluding the human time for approvals and purchased game models. The estimate for people is based on the usual development time of a 3D game of similar scope with a team of seven. Today the agents are mostly slowed down by the limits of the shared CI.

      Specialized sessions

      Each session has its own domain and its own label in the task queue.

      • Graphics Shaders, lighting, effects and rendering performance.
      • Game systems Game rules, inventory, combat, saving state.
      • World and economy Trade, resources, prices and the balance of the game world.
      • Terrain Landscape generation, height data, transitions and navigation.
      • Models Preparing and optimizing 3D models and loading them.
      • Tester Automated playthroughs, screenshots, bug reports back into the queue.
      task queue
      • Terrain Smoother transition between terrain tiles status: in progress
      • Graphics Shadows on distant objects review
      • Tester Regression: saving the game after loading a map status: new
      • Decision Approve merge into main awaiting approval
      Illustrative example of the queue. The label decides which session takes the task.

      Rules of collaboration

      1. Tasks as GitHub Issues The dispatcher creates and triages issues. The label decides which session takes the task. Discussion and results stay with the task.
      2. Shared rules in AGENTS.md One file in the repository describes conventions, domain boundaries and the workflow. Every session reads it before starting a task.
      3. A separate branch for each session Sessions never work on the same branch. Changes only meet in a pull request, where it is clear what changes and why.
      4. CI and deployment After a merge into main, CI builds the game and deploys it to Cloudflare. Nobody uploads files by hand.
      night

      Agents work independently on prepared tasks. The result: pull requests, test results and new bugs in the queue.

      day

      A human goes through the results, approves or returns changes, adjusts priorities and settles disputed points.

      Tools and small jobs

      Tools

      What it runs on

      • Claude Code

        agents

        An agent in the terminal that reads the repository, runs commands and writes code. The individual roles run on it.

      • GitHub Issues + Actions

        queue and CI

        Issues serve as a labelled task queue; Actions run the build, tests and deployment after a merge.

      • Playwright

        tests

        An automated browser in which the Tester walks through the application and takes screenshots for comparison.

      • Cloudflare Pages, R2, Workers

        hosting

        Pages hosts static builds, R2 stores larger files and Workers handle server-side logic.

      • Azure

        alternative

        An option for companies with infrastructure and agreements at Microsoft: VMs for agents, hosting and identity management.

      Small jobs

      Hosting and domain setup at the same hourly rate as a regular supplier.

      • Operations and infrastructure over 5×faster

        Website deployment on Cloudflare Pages

        Connecting the repository, automated builds, branch previews, HTTPS and security headers.

        DevOps specialist at a supplier 4 hours
        NetVoice with agents 45 minutes
        Supplier≈ 6,000 CZK NetVoice≈ 1,100 CZK

        80–90% cheaper you save ≈ 4,900 CZK

        Of which 15 minutes of human work (review and approval).

      • Operations and infrastructure 2.6×faster

        DNS at Forpsi

        Pointing the domain to the website plus email records (SPF, DKIM, DMARC). The agent prepares the exact records, a human enters them in the Forpsi console.

        DevOps specialist at a supplier 2 hours
        NetVoice with agents 45 minutes
        Supplier≈ 3,000 CZK NetVoice≈ 1,100 CZK

        62% cheaper you save ≈ 1,900 CZK

        Of which 15 minutes of human work (entering the records in the Forpsi console). DNS propagation (up to 24 hours) takes the same time either way.