Yes. Now that the architecture and assumptions are clear, I would build this in phases, with each phase producing something usable and measurable. We should not jump directly to cloud auto-spawning.

Alpha Search — Master/Slave Auto-Scaling Roadmap

Target architecture

Ultimately we want:

                           USERS
                             │
                             ▼
                    ┌────────────────┐
                    │     MASTER     │
                    │                │
                    │ Traffic Router │
                    │ Search Control │
                    │ Slave Manager  │
                    │ Result Merger  │
                    └───────┬────────┘
                            │
              ┌─────────────┼─────────────┐
              │             │             │
              ▼             ▼             ▼
          ┌────────┐    ┌────────┐    ┌────────┐
          │ SLAVE 1│    │ SLAVE 2│    │ SLAVE 3│
          │ IP-A   │    │ IP-B   │    │ IP-C   │
          └───┬────┘    └───┬────┘    └───┬────┘
              │              │              │
              ▼              ▼              ▼
           SearXNG        SearXNG        SearXNG
              │              │              │
              ▼              ▼              ▼
          Crawl4AI        Crawl4AI        Crawl4AI

With an observability/control plane watching everything:

                    ┌─────────────────────┐
                    │    OBSERVABILITY    │
                    │                     │
                    │ Traffic             │
                    │ Search quality      │
                    │ Providers           │
                    │ Crawling            │
                    │ Slave lifecycle     │
                    │ Costs               │
                    └──────────┬──────────┘
                               │
                               ▼
                         Control decisions

Phase 0 — Freeze the current foundation

Status: Mostly complete ✅

Before introducing distributed architecture, establish the current baseline.

We already have:

Search discovery

Alpha
 ↓
ONE SearXNG request
 ↓
N configured engines

Search admission

20 active searches
+
unlimited waiting

Crawl admission

CRAWL_GLOBAL_CONCURRENCY=10

Search pipeline

SearXNG
 ↓
dedup
 ↓
lexical ranking
 ↓
TOP_N
 ↓
Crawl4AI

Configuration

Everything important should remain configuration-driven.

Outcome: stable single-instance Alpha Search.


Phase 1 — Define the Master/Slave contract

Next major phase

Before spawning anything, define exactly what a slave is.

A slave should expose a small internal API such as:

POST /internal/search
GET  /internal/health
GET  /internal/status

Potential search request:

{
  "requestId": "uuid",
  "query": "Flutter Bloc architecture",
  "searchProfile": "default"
}

Potential response:

{
  "requestId": "uuid",
  "results": [...],
  "metadata": {
    "durationMs": 2410,
    "providers": [...],
    "crawlSuccessCount": 5
  }
}

The important thing is that Master doesn't need to know how the slave performs the search.

It only knows:

send search
→ receive result

Also define slave identity

Every slave should have:

slaveId
instanceId
createdAt
ip/endpoint
status
configuration/profile
lastHeartbeat

Slave states

PROVISIONING
     ↓
STARTING
     ↓
READY
     ↓
BUSY
     ↓
DRAINING
     ↓
TERMINATING
     ↓
TERMINATED

Deliverable: a documented and implemented Master ↔ Slave protocol.


Phase 2 — Separate Master and Search Worker responsibilities

Currently Alpha Search is essentially one application.

We should logically separate:

MASTER
 ├── incoming traffic
 ├── routing
 ├── slave registry
 ├── lifecycle decisions
 └── result coordination

SLAVE
 ├── SearchService
 ├── SearXNG
 ├── ranking
 ├── Crawl4AI
 └── result generation

Don't necessarily create two repositories yet.

We can first create clear application/module boundaries.

This is important because eventually:

Master

should be lightweight while:

Slave

is where the expensive search execution happens.

Deliverable: Alpha Search can conceptually operate in master/worker mode without cloud provisioning yet.


Phase 3 — Local slave simulation

Before touching AWS/Hetzner/DigitalOcean/etc., simulate slaves locally.

For example:

Master : 3000

Slave 1 : 3101
Slave 2 : 3102
Slave 3 : 3103

Master discovers:

slave-1 → localhost:3101
slave-2 → localhost:3102
slave-3 → localhost:3103

Then test:

User
 ↓
Master
 ↓
Slave
 ↓
Search
 ↓
Master
 ↓
User

Why this phase is critical

We can solve:

without paying for cloud servers.

Deliverable: distributed architecture works on one machine.


Phase 4 — Master traffic routing

Now implement the actual routing decision.

Initially, keep it simple:

Available slaves
       ↓
Choose least-loaded READY slave
       ↓
Forward request

Each slave reports:

activeRequests
capacity
status

Example:

Slave 1: 17/20
Slave 2:  4/20
Slave 3: 11/20

Master chooses:

Slave 2

This is better than blindly round-robin because our search workload is not uniform.

Important

The existing slave-level:

SEARCH_CONCURRENCY=20

still remains.

Master routing simply determines which slave receives the request.


Phase 5 — Shared state / coordination

Now we reach the point where Redis becomes useful.

Something like:

                 Redis
                   │
       ┌───────────┼───────────┐
       ▼           ▼           ▼
    Master      Slave 1     Slave 2

Redis can hold/coordinate:

slave registry
slave heartbeat
active workload
slave status
distributed locks
lifecycle state

Potentially later:

distributed queues
leases
scaling decisions

Important

Don't introduce Redis before we actually need distributed coordination.

Deliverable: multiple Master/Slave processes can share consistent state.


Phase 6 — Slave provisioning

Now connect the system to your actual server provider.

Create a Slave Provisioner abstraction:

SlaveProvisioner
       │
       ├── create()
       ├── waitUntilReady()
       ├── getStatus()
       ├── drain()
       └── terminate()

The Alpha Search code shouldn't care whether the infrastructure is:

The infrastructure adapter handles that.

Provisioning flow

Master decides:
"Need slave"

       ↓

Provisioner.create()

       ↓

Cloud server created

       ↓

Bootstrap Alpha Slave

       ↓

Slave starts

       ↓

Health check

       ↓

Slave registers

       ↓

READY

Only after READY should traffic be routed to it.


Phase 7 — Traffic-based autoscaling

Now implement your first scaling rule.

Initial policy:

TARGET_REQUESTS_PER_MINUTE = 15
MAX_SLAVES = configurable

For example:

1 slave
15 req/min target

traffic = 22 req/min

→ capacity insufficient
→ provision Slave 2

Then:

2 slaves
≈ 30 req/min target

But remember:

15 is a starting policy, not a physical guarantee.

The observability system will eventually tell us whether it should become:

10
15
20
30
...

Phase 8 — Search-quality scaling

Now implement your second scaling signal.

For every search:

Search
 ↓
provider results
 ↓
dedup
 ↓
ranking
 ↓
quality evaluation

We need a deterministic definition of:

Enough good results

Initially that can be:

goodResults >= TOP_N_RESULTS_TO_CRAWL

For example:

TOP_N = 5

Good results = 5
→ success

But:

Good results = 2
→ insufficient

Then Master can decide whether to execute a secondary search strategy.

This is where your different-slave configuration becomes useful.

For example:

Primary Slave
Google + Yep + DDG + Mojeek
       ↓
insufficient
       ↓
Secondary Slave
different egress IP
different provider profile
       ↓
results
       ↓
merge
       ↓
dedup
       ↓
rank
       ↓
TOP_N

This lets us experiment with your IP hypothesis without assuming it works beforehand.


Phase 9 — Search profiles

This is an important prerequisite for Phase 8.

Instead of hardcoding:

every slave = same providers

introduce a concept like:

SearchProfile A
Google + Yep + DDG + Mojeek

SearchProfile B
Yep + Brave API + another provider

SearchProfile C
different provider combination

Then a slave can be provisioned with:

profile=A

or:

profile=B

This makes the search-quality fallback strategy configurable rather than embedding provider decisions into scaling code.


Phase 10 — Slave lifecycle management

Now implement:

Maximum slaves

MAX_SLAVES=5

Never exceed it.

Lifetime

SLAVE_MAX_LIFETIME=2h

At expiry:

READY
 ↓
DRAINING
 ↓
stop accepting new requests
 ↓
finish active requests
 ↓
terminate

Failed termination

TERMINATING
 ↓
termination failed
 ↓
FAILED_TERMINATION
 ↓
alert

The Master must retain enough state to identify the server for manual cleanup.


Phase 11 — Observability system

This should become a proper subsystem rather than scattered logs.

Master metrics

requests/min
active requests
queued requests
latency
errors

Slave metrics

CPU
memory
requests/min
active requests
latency
uptime
health
age

Provider metrics

requests
success
429
403
CAPTCHA
timeouts
latency
success rate

Search quality

results returned
results after dedup
good results
TOP_N
provider contribution

Lifecycle

spawn requested
spawn started
spawn ready
spawn failed
drain started
termination started
termination completed
termination failed

Phase 12 — Dashboard

Now expose this operationally.

Something like:

┌─────────────────────────────────────────────┐
│ ALPHA SEARCH                                │
├─────────────────────────────────────────────┤
│ Requests/min       27                       │
│ Active searches    34                       │
│ Slaves             3 / 5                    │
│                                             │
│ SEARCH QUALITY                              │
│ Good result rate  96.2%                     │
│                                             │
│ SLAVES                                      │
│                                             │
│ slave-01   READY      11/20   42m           │
│ slave-02   READY       8/20   31m           │
│ slave-03   DRAINING    3/20   2h            │
│                                             │
│ PROVIDERS                                   │
│ Google       91% ✓                          │
│ Yep          99% ✓                          │
│ DDG          78% ⚠                          │
│ Mojeek       84% ✓                          │
└─────────────────────────────────────────────┘

This is where you'll actually be able to make informed decisions.


Phase 13 — Dynamic/adaptive scaling

Only after all the above is working.

Then we can replace simple rules with a proper control loop:

                    Metrics
                       │
          ┌────────────┼────────────┐
          ▼            ▼            ▼
       Traffic      Quality       Cost
          │            │            │
          └────────────┼────────────┘
                       ▼
                 Scaling Engine
                       │
              ┌────────┴────────┐
              ▼                 ▼
            SCALE UP          SCALE DOWN
              │                 │
          provision           drain
              │                 │
            ready            terminate

Eventually it could dynamically determine:

Current:
3 slaves

Traffic:
41 req/min

Quality:
98%

→ keep 3

versus:

Traffic:
72 req/min

Quality:
99%

→ spawn 2

versus:

Traffic:
12 req/min

Quality:
99%

Slave 3:
mostly idle

→ drain Slave 3

Phase 14 — Production hardening

Finally:


The execution order I'd use

This is important because you want AG tasks small and trackable.

PHASE 0   Current single-instance foundation       ✅
    ↓
PHASE 1   Master ↔ Slave contract
    ↓
PHASE 2   Separate master/worker responsibilities
    ↓
PHASE 3   Local slave simulation
    ↓
PHASE 4   Master routing
    ↓
PHASE 5   Redis/shared coordination
    ↓
PHASE 6   Cloud slave provisioning
    ↓
PHASE 7   Traffic-based autoscaling
    ↓
PHASE 8   Search-quality-based scaling
    ↓
PHASE 9   Search profiles
    ↓
PHASE 10  Slave lifecycle management
    ↓
PHASE 11  Observability
    ↓
PHASE 12  Dashboard
    ↓
PHASE 13  Dynamic/adaptive scaling
    ↓
PHASE 14  Production hardening

One adjustment I'd make to the ordering

I would actually bring Search Profiles (Phase 9) slightly earlier, before implementing search-quality autoscaling.

Because the quality scaler needs to know:

"What different search execution am I asking the new slave to perform?"

So the practical sequence becomes:

Master/Slave contract
        ↓
Local slaves
        ↓
Routing
        ↓
Search profiles
        ↓
Shared coordination
        ↓
Cloud provisioning
        ↓
Traffic autoscaling
        ↓
Quality-based secondary search
        ↓
Lifecycle management
        ↓
Observability/dashboard
        ↓
Adaptive scaling

Most important principle

Do not build "the autoscaler" as one giant feature.

We're essentially building a small distributed platform:

              ┌────────────────────┐
              │      MASTER        │
              │                    │
              │ Routing            │
              │ Scaling            │
              │ Lifecycle          │
              │ Aggregation        │
              └─────────┬──────────┘
                        │
               ┌────────┴────────┐
               ▼                 ▼
          Coordination       Provisioning
             Redis             Cloud API
               │                 │
               └────────┬────────┘
                        │
                   ┌────┴────┐
                   ▼         ▼
                Slaves     Metrics
                             │
                             ▼
                          Dashboard

So we'll take one box at a time.

Immediate next task

I would not give AG a cloud/autoscaling task yet.

The first task should be:

Design and implement the Master ↔ Slave internal contract and basic Slave identity/health model, while keeping everything local.

Once that works, we can simulate two or three slaves locally and prove the architecture before spending money on actual server instances.