Yes. Now that the architecture and assumptions are clear, I would build this in phases, with each phase producing something usable and measurable. We should not jump directly to cloud auto-spawning.
Ultimately we want:
USERS
│
▼
┌────────────────┐
│ MASTER │
│ │
│ Traffic Router │
│ Search Control │
│ Slave Manager │
│ Result Merger │
└───────┬────────┘
│
┌─────────────┼─────────────┐
│ │ │
▼ ▼ ▼
┌────────┐ ┌────────┐ ┌────────┐
│ SLAVE 1│ │ SLAVE 2│ │ SLAVE 3│
│ IP-A │ │ IP-B │ │ IP-C │
└───┬────┘ └───┬────┘ └───┬────┘
│ │ │
▼ ▼ ▼
SearXNG SearXNG SearXNG
│ │ │
▼ ▼ ▼
Crawl4AI Crawl4AI Crawl4AI
With an observability/control plane watching everything:
┌─────────────────────┐
│ OBSERVABILITY │
│ │
│ Traffic │
│ Search quality │
│ Providers │
│ Crawling │
│ Slave lifecycle │
│ Costs │
└──────────┬──────────┘
│
▼
Control decisions
Status: Mostly complete ✅
Before introducing distributed architecture, establish the current baseline.
We already have:
Alpha
↓
ONE SearXNG request
↓
N configured engines
20 active searches
+
unlimited waiting
CRAWL_GLOBAL_CONCURRENCY=10
SearXNG
↓
dedup
↓
lexical ranking
↓
TOP_N
↓
Crawl4AI
Everything important should remain configuration-driven.
Outcome: stable single-instance Alpha Search.
Next major phase
Before spawning anything, define exactly what a slave is.
A slave should expose a small internal API such as:
POST /internal/search
GET /internal/health
GET /internal/status
Potential search request:
{
"requestId": "uuid",
"query": "Flutter Bloc architecture",
"searchProfile": "default"
}
Potential response:
{
"requestId": "uuid",
"results": [...],
"metadata": {
"durationMs": 2410,
"providers": [...],
"crawlSuccessCount": 5
}
}
The important thing is that Master doesn't need to know how the slave performs the search.
It only knows:
send search
→ receive result
Every slave should have:
slaveId
instanceId
createdAt
ip/endpoint
status
configuration/profile
lastHeartbeat
PROVISIONING
↓
STARTING
↓
READY
↓
BUSY
↓
DRAINING
↓
TERMINATING
↓
TERMINATED
Deliverable: a documented and implemented Master ↔ Slave protocol.
Currently Alpha Search is essentially one application.
We should logically separate:
MASTER
├── incoming traffic
├── routing
├── slave registry
├── lifecycle decisions
└── result coordination
SLAVE
├── SearchService
├── SearXNG
├── ranking
├── Crawl4AI
└── result generation
Don't necessarily create two repositories yet.
We can first create clear application/module boundaries.
This is important because eventually:
Master
should be lightweight while:
Slave
is where the expensive search execution happens.
Deliverable: Alpha Search can conceptually operate in master/worker mode without cloud provisioning yet.
Before touching AWS/Hetzner/DigitalOcean/etc., simulate slaves locally.
For example:
Master : 3000
Slave 1 : 3101
Slave 2 : 3102
Slave 3 : 3103
Master discovers:
slave-1 → localhost:3101
slave-2 → localhost:3102
slave-3 → localhost:3103
Then test:
User
↓
Master
↓
Slave
↓
Search
↓
Master
↓
User
We can solve:
without paying for cloud servers.
Deliverable: distributed architecture works on one machine.
Now implement the actual routing decision.
Initially, keep it simple:
Available slaves
↓
Choose least-loaded READY slave
↓
Forward request
Each slave reports:
activeRequests
capacity
status
Example:
Slave 1: 17/20
Slave 2: 4/20
Slave 3: 11/20
Master chooses:
Slave 2
This is better than blindly round-robin because our search workload is not uniform.
The existing slave-level:
SEARCH_CONCURRENCY=20
still remains.
Master routing simply determines which slave receives the request.
Now we reach the point where Redis becomes useful.
Something like:
Redis
│
┌───────────┼───────────┐
▼ ▼ ▼
Master Slave 1 Slave 2
Redis can hold/coordinate:
slave registry
slave heartbeat
active workload
slave status
distributed locks
lifecycle state
Potentially later:
distributed queues
leases
scaling decisions
Don't introduce Redis before we actually need distributed coordination.
Deliverable: multiple Master/Slave processes can share consistent state.
Now connect the system to your actual server provider.
Create a Slave Provisioner abstraction:
SlaveProvisioner
│
├── create()
├── waitUntilReady()
├── getStatus()
├── drain()
└── terminate()
The Alpha Search code shouldn't care whether the infrastructure is:
The infrastructure adapter handles that.
Master decides:
"Need slave"
↓
Provisioner.create()
↓
Cloud server created
↓
Bootstrap Alpha Slave
↓
Slave starts
↓
Health check
↓
Slave registers
↓
READY
Only after READY should traffic be routed to it.
Now implement your first scaling rule.
Initial policy:
TARGET_REQUESTS_PER_MINUTE = 15
MAX_SLAVES = configurable
For example:
1 slave
15 req/min target
traffic = 22 req/min
→ capacity insufficient
→ provision Slave 2
Then:
2 slaves
≈ 30 req/min target
But remember:
15 is a starting policy, not a physical guarantee.
The observability system will eventually tell us whether it should become:
10
15
20
30
...
Now implement your second scaling signal.
For every search:
Search
↓
provider results
↓
dedup
↓
ranking
↓
quality evaluation
We need a deterministic definition of:
Enough good results
Initially that can be:
goodResults >= TOP_N_RESULTS_TO_CRAWL
For example:
TOP_N = 5
Good results = 5
→ success
But:
Good results = 2
→ insufficient
Then Master can decide whether to execute a secondary search strategy.
This is where your different-slave configuration becomes useful.
For example:
Primary Slave
Google + Yep + DDG + Mojeek
↓
insufficient
↓
Secondary Slave
different egress IP
different provider profile
↓
results
↓
merge
↓
dedup
↓
rank
↓
TOP_N
This lets us experiment with your IP hypothesis without assuming it works beforehand.
This is an important prerequisite for Phase 8.
Instead of hardcoding:
every slave = same providers
introduce a concept like:
SearchProfile A
Google + Yep + DDG + Mojeek
SearchProfile B
Yep + Brave API + another provider
SearchProfile C
different provider combination
Then a slave can be provisioned with:
profile=A
or:
profile=B
This makes the search-quality fallback strategy configurable rather than embedding provider decisions into scaling code.
Now implement:
MAX_SLAVES=5
Never exceed it.
SLAVE_MAX_LIFETIME=2h
At expiry:
READY
↓
DRAINING
↓
stop accepting new requests
↓
finish active requests
↓
terminate
TERMINATING
↓
termination failed
↓
FAILED_TERMINATION
↓
alert
The Master must retain enough state to identify the server for manual cleanup.
This should become a proper subsystem rather than scattered logs.
requests/min
active requests
queued requests
latency
errors
CPU
memory
requests/min
active requests
latency
uptime
health
age
requests
success
429
403
CAPTCHA
timeouts
latency
success rate
results returned
results after dedup
good results
TOP_N
provider contribution
spawn requested
spawn started
spawn ready
spawn failed
drain started
termination started
termination completed
termination failed
Now expose this operationally.
Something like:
┌─────────────────────────────────────────────┐
│ ALPHA SEARCH │
├─────────────────────────────────────────────┤
│ Requests/min 27 │
│ Active searches 34 │
│ Slaves 3 / 5 │
│ │
│ SEARCH QUALITY │
│ Good result rate 96.2% │
│ │
│ SLAVES │
│ │
│ slave-01 READY 11/20 42m │
│ slave-02 READY 8/20 31m │
│ slave-03 DRAINING 3/20 2h │
│ │
│ PROVIDERS │
│ Google 91% ✓ │
│ Yep 99% ✓ │
│ DDG 78% ⚠ │
│ Mojeek 84% ✓ │
└─────────────────────────────────────────────┘
This is where you'll actually be able to make informed decisions.
Only after all the above is working.
Then we can replace simple rules with a proper control loop:
Metrics
│
┌────────────┼────────────┐
▼ ▼ ▼
Traffic Quality Cost
│ │ │
└────────────┼────────────┘
▼
Scaling Engine
│
┌────────┴────────┐
▼ ▼
SCALE UP SCALE DOWN
│ │
provision drain
│ │
ready terminate
Eventually it could dynamically determine:
Current:
3 slaves
Traffic:
41 req/min
Quality:
98%
→ keep 3
versus:
Traffic:
72 req/min
Quality:
99%
→ spawn 2
versus:
Traffic:
12 req/min
Quality:
99%
Slave 3:
mostly idle
→ drain Slave 3
Finally:
This is important because you want AG tasks small and trackable.
PHASE 0 Current single-instance foundation ✅
↓
PHASE 1 Master ↔ Slave contract
↓
PHASE 2 Separate master/worker responsibilities
↓
PHASE 3 Local slave simulation
↓
PHASE 4 Master routing
↓
PHASE 5 Redis/shared coordination
↓
PHASE 6 Cloud slave provisioning
↓
PHASE 7 Traffic-based autoscaling
↓
PHASE 8 Search-quality-based scaling
↓
PHASE 9 Search profiles
↓
PHASE 10 Slave lifecycle management
↓
PHASE 11 Observability
↓
PHASE 12 Dashboard
↓
PHASE 13 Dynamic/adaptive scaling
↓
PHASE 14 Production hardening
I would actually bring Search Profiles (Phase 9) slightly earlier, before implementing search-quality autoscaling.
Because the quality scaler needs to know:
"What different search execution am I asking the new slave to perform?"
So the practical sequence becomes:
Master/Slave contract
↓
Local slaves
↓
Routing
↓
Search profiles
↓
Shared coordination
↓
Cloud provisioning
↓
Traffic autoscaling
↓
Quality-based secondary search
↓
Lifecycle management
↓
Observability/dashboard
↓
Adaptive scaling
Do not build "the autoscaler" as one giant feature.
We're essentially building a small distributed platform:
┌────────────────────┐
│ MASTER │
│ │
│ Routing │
│ Scaling │
│ Lifecycle │
│ Aggregation │
└─────────┬──────────┘
│
┌────────┴────────┐
▼ ▼
Coordination Provisioning
Redis Cloud API
│ │
└────────┬────────┘
│
┌────┴────┐
▼ ▼
Slaves Metrics
│
▼
Dashboard
So we'll take one box at a time.
I would not give AG a cloud/autoscaling task yet.
The first task should be:
Design and implement the Master ↔ Slave internal contract and basic Slave identity/health model, while keeping everything local.
Once that works, we can simulate two or three slaves locally and prove the architecture before spending money on actual server instances.