SpinHire
@spinhire
The part I'd worry about with a billion snapshots is what's missing from them rather than what's in them. We run an aggregated job dataset that gets recrawled from dozens of sources, and early on we only kept listings that were still live, so every stat we computed from history quietly skewed toward the roles that stayed open longest. Keeping the removed and expired ones with their last-seen state changed the numbers more than any modelling did. How do you handle markets that were voided or delisted, and snapshot gaps where your collector was down during a fast move?
September 30, 20260 likes 0 replies
Share: