# Market Intel Scrape — Timeout Fix (2026-09-28)

## The problem

`RunCompetitorScrapeJob` had `$timeout = 3600` (1 hour) — Laravel's queue worker hard-kills a job at its own
timeout. Two competitors (Its, Maltings) genuinely need longer than an hour to get through their per-run URL
cap, so the worker was killing them mid-run every night, showing up in the console as red
`Unconfirmed · <competitor> · Amazon rejected this after initial acceptance` — style failures (actually a
generic "has timed out" kill, not a real scrape error) and leaving `ScrapeRun` rows stuck.

## The fix

`plugins/webkul/market-intel/src/Jobs/RunCompetitorScrapeJob.php`:

- **Hard timeout**: 3600 → **7800** (2h10m). Must stay under the database queue's `retry_after`
  (`config/queue.php`, 10800s / 3h default) or the job could be picked up a second time while still running.
- **Soft time budget**: `TIME_BUDGET_SECONDS = 7200` (2h). `scrapeUrls()` now checks elapsed time on every
  iteration and, once past it, stops cleanly and returns — same as the existing cancel-request check just above
  it — instead of letting the hard timeout kill it. The run is recorded as **Completed**, not failed, with a
  `stoppedAtTimeBudget` flag; the job console line reads e.g.
  `Its · 2,400 scraped of 6,000 (stopped at the 2h time limit, next run continues)`.
- **Why stopping early is safe**: `DiscoveredUrl`s are walked oldest-scraped-first
  (`orderByRaw('last_scraped_at IS NOT NULL, last_scraped_at ASC')`), so the next scheduled run picks up exactly
  where this one left off — nothing is skipped permanently, it just spreads across more runs.

## Deploy note

Restart the queue workers after deploying so they pick up the new job timeout — a running worker process keeps
whatever timeout it read at boot: `sudo supervisorctl restart cydekick-worker:*`.

## Still true after this fix

The scrape being slow in the first place isn't addressed — this only stops the "took too long" case counting as
a failure. If Its/Maltings need to finish faster, that's the competitor's `max_urls_per_run` / delay-between-
requests settings (`Competitor::getMaxUrlsPerRun()`, `getScrapeDelayMs()`), not this job.
