Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
15 changes: 15 additions & 0 deletions .github/workflows/canary.yml
Original file line number Diff line number Diff line change
Expand Up @@ -119,6 +119,18 @@ jobs:
f"the {floor:,} floor")
if len(r.get("sku") or "") != 36:
fail.append(f"@{name}: sku {r.get('sku')!r} is not a UUID")
# Snapchat's page state is not a published contract. A served page
# that has lost a key the parser reads gives silent nulls, so the
# parser names such keys and this fails on them by name.
for label, m in (("profiles", meta),):
if m.get("payload_keys_missing"):
fail.append(f"{label}: Snapchat moved key(s) the parser "
f"reads: {m['payload_keys_missing']}")
# The default transport is plain HTTP; a browser here means the
# site refused it, which is the README's central claim failing.
if meta.get("transport") != "http":
fail.append(f"transport {meta.get('transport')!r}: the default "
f"run did not stay on plain HTTP")
espn = by.get("espn")
if espn is None or espn.get("public_profile") is not False:
fail.append(f"@espn: expected an ordinary account row, got "
Expand All @@ -140,6 +152,9 @@ jobs:
if bad:
fail.append(f"spotlight: {len(bad)} row(s) without a positive "
f"view_count — an empty slot read as a video?")
if smeta.get("payload_keys_missing"):
fail.append(f"spotlight: Snapchat moved key(s) the parser "
f"reads: {smeta['payload_keys_missing']}")
pairs = [(r["page"], r["position"]) for r in spot]
if len(set(pairs)) != len(pairs):
fail.append("spotlight: page+position is not unique")
Expand Down
2 changes: 1 addition & 1 deletion .github/workflows/tests.yml
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,7 @@ jobs:
matrix:
# 3.9 is the floor the README claims. Claiming it without testing it is
# how a walrus operator or a `X | None` annotation ships and breaks it.
python-version: ["3.9", "3.12"]
python-version: ["3.9", "3.14"]

steps:
- uses: actions/checkout@v4
Expand Down
39 changes: 39 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,45 @@ closely as a CLI toolkit can. A patch release means **fixes** — it does not
promise that every flag's default is frozen, and where a default does change
in one, the note leads with it.

## [0.2.0] — 2026-10-05

Changes from a third-party audit, each checked against a live run before it
was accepted. Two columns are added to the profile row, which is why this
is a minor release; no column was removed or renamed.

### Added

- **`has_more_spotlight` and `has_more_highlights` on the profile row**,
from the page's own cursors. The page lists about 25 Spotlight videos
and a slice of highlights; these say when the account has more, on the
row itself, so a row read without its sidecar still says its counts are
a slice.
- **`payload_keys_missing` in the sidecar.** The parser now checks that
every `__NEXT_DATA__` key it reads is present, and names the ones that
are not, per account. A page that renders but has lost a field used to
produce silent nulls. Zero false positives on 20 raw captures and a live
page; the daily canary fails on a non-empty value.
- **`transport` records what actually fetched the pages** (`http`,
`browser`, `cdp`), with `transport_requested` beside it. It used to
record the flag, so the default run said `auto` and never which one.
- **CSV formula neutralisation**, lifted from rakuten-scraper: a string cell
beginning `=`, `+`, `-`, `@` or a control character is prefixed with `'`
in CSV only, and counted in `csv_cells_escaped`. It does not fire on
today's data (0 of 10,554 string cells, measured); it is here because
the text is the account owner's.
- **A check that the three engines' shared code is identical**, so the
three copies cannot drift apart unnoticed.
- CI tests the newest end of the supported range on Python 3.14 (was 3.12).

### Fixed

- **Output files were created owner-only (0600).** The atomic writer's
temporary file is 0600 and the rename kept it: nine of nine files on a
live run under umask 022. New files now get the umask's mode (0644
there). A file that already exists keeps its mode, so outputs written by
0.1.0 stay 0600 until you delete them once — the writer cannot tell a
mode someone chose from one the bug left.

## [0.1.0] — 2026-09-24

First release.
Expand Down
27 changes: 22 additions & 5 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@
[![release](https://img.shields.io/github/v/release/2scraper/snapchat-scraper?sort=semver)](https://github.com/2scraper/snapchat-scraper/releases)
[![tests](https://github.com/2scraper/snapchat-scraper/actions/workflows/tests.yml/badge.svg)](https://github.com/2scraper/snapchat-scraper/actions/workflows/tests.yml)
[![canary](https://github.com/2scraper/snapchat-scraper/actions/workflows/canary.yml/badge.svg)](https://github.com/2scraper/snapchat-scraper/actions/workflows/canary.yml)
[![python](https://img.shields.io/badge/python-3.9%20%7C%203.12-blue)](pyproject.toml)
[![python](https://img.shields.io/badge/python-3.9%20%7C%203.14-blue)](pyproject.toml)
[![licence](https://img.shields.io/badge/licence-MIT-green)](LICENSE)
[![engines](https://img.shields.io/badge/engines-Playwright%20%7C%20Selenium%20%7C%20pyppeteer%20%7C%20CDP-informational)](#engines-and-what-each-one-costs-you)
[![runs without an account](https://img.shields.io/badge/runs%20without-an%20account-brightgreen)](#you-do-not-need-a-key-a-proxy-or-an-account)
Expand Down Expand Up @@ -130,9 +130,10 @@ Measured on 2026-09-24: @nasa — 757,800 subscribers, 7 Spotlight videos,
* **At most ~25 Spotlight videos per account.** That is what a profile
page lists. The page carries a cursor for the rest, consumed by
Snapchat's own protobuf API, which **this repo does not implement**. The
sidecar names every account whose page said there was more, in
`handles_with_more_than_page`, so "complete" is never misread as "the
account's whole history".
profile row says so itself, in `has_more_spotlight` and
`has_more_highlights`, and the sidecar names every such account in
`handles_with_more_than_page`. `status: complete` means every account
asked for was answered — never "the account's whole history".
* **An account with one row and nothing in it.** An ordinary account —
not a Public Profile — is served as a username and a Snapcode and
nothing else (@espn, measured). `--mode profile` emits a row with
Expand Down Expand Up @@ -230,7 +231,23 @@ reason, **which** accounts failed by number, `handles_unavailable`,
`handles_without_public_profile`, `handles_with_more_than_page`,
`spotlight_empty_slots`, and `viewer_countries` — the country Snapchat
says the request came from, which is how you check that a proxy's or a
Scraping Browser's `country-` segment did what you asked.
Scraping Browser's `country-` segment did what you asked. Also:

* `transport` — what actually fetched the pages (`http`, `browser` or
`cdp`), beside `transport_requested`. The default `auto` is HTTP until
the site refuses, so this is the field that says whether a browser was
ever involved; `engine` names the script that ran.
* `payload_keys_missing` — keys this parser reads that a served page no
longer carried, with the accounts each was missing on. The source is
Snapchat's own page state, not a published contract, so a page can
render and still have moved a field; this is how that shows up instead
of as silent nulls. Empty on every page this repo was tested against,
and the daily canary fails if it is not.
* `csv_cells_escaped` — CSV cells that began with `=`, `+`, `-`, `@` or a
control character and were prefixed with `'` so a spreadsheet reads
them as text (bios and captions are written by account owners). JSON
keeps the bytes as served. 0 of 10,554 string cells on a live run of
2026-10-05.

---

Expand Down
101 changes: 96 additions & 5 deletions output_writer.py
Original file line number Diff line number Diff line change
Expand Up @@ -137,6 +137,13 @@ class Profile:
has_story: Optional[bool] = None
has_curated_highlights: Optional[bool] = None
has_spotlight_highlights: Optional[bool] = None
# True when the page carried a cursor for MORE Spotlight videos /
# highlights than it listed. The counts above are then the page's
# slice, not the account's total. On the ROW rather than only in the
# sidecar, because a row that reaches a consumer without its sidecar
# must still say so (third-party audit, 2026-10-05).
has_more_spotlight: Optional[bool] = None
has_more_highlights: Optional[bool] = None

# ---- what the account says about itself --------------------------------
bio: Optional[str] = None
Expand Down Expand Up @@ -314,6 +321,28 @@ def _csv_value(v: Any) -> Any:
return v


def _target_mode(path: str) -> int:
"""The permission bits the finished file should carry.

`NamedTemporaryFile` creates its file 0600 and `os.replace` keeps the
mode, so every output this module wrote was readable by its owner alone
— measured 2026-10-05 on a live run under umask 022: nine of nine
`.json`/`.csv`/`.meta.json` files came out 0600, where the `open(path,
"w")` this writer replaced would have given 0644. A file a cron job
writes and another user's pipeline reads then fails on permission
rather than on content.

An EXISTING target keeps its mode, because someone may have tightened
it on purpose; a new one gets what `open()` would have given it.
"""
try:
return os.stat(path).st_mode & 0o777
except OSError:
umask = os.umask(0)
os.umask(umask)
return 0o666 & ~umask


@contextlib.contextmanager
def _atomic(path: str, newline: Optional[str] = None):
"""Write to a temporary file beside `path`, then rename over it.
Expand Down Expand Up @@ -345,6 +374,7 @@ def _atomic(path: str, newline: Optional[str] = None):
yield handle
handle.flush()
os.fsync(handle.fileno())
os.chmod(handle.name, _target_mode(path))
os.replace(handle.name, path)
except BaseException:
# Leave the destination untouched. A failed write must not be
Expand All @@ -361,7 +391,42 @@ def write_json(rows: Sequence[Any], path: str) -> None:
json.dump([asdict(r) for r in rows], f, ensure_ascii=False, indent=2)


def write_csv(rows: Sequence[Any], path: str, row_cls: Type = Profile) -> None:
# A spreadsheet treats a cell beginning with one of these as a FORMULA, not
# as text, and the bio, title, address and caption columns here are written
# by whoever owns the account. `=HYPERLINK(...)`, `+cmd|...` and `@SUM(...)`
# are the classic shapes; the tab and the newline are here because a leading
# one is stripped by some readers, which exposes whatever follows it.
#
# Measured on this site on 2026-10-05 before shipping the guard: 0 of 10,554
# string cells across a live run of five accounts in all three modes begin
# with any of them. So this does not fire on today's data and is not claimed
# to; it is here because the text is not ours. Lifted verbatim from
# rakuten-scraper, where it was written first.
CSV_FORMULA_LEADS = ("=", "+", "-", "@", "\t", "\r", "\n")

# Prefixed to a cell that would otherwise be read as a formula. Excel, Google
# Sheets and LibreOffice all treat the apostrophe as "the rest of this cell
# is text" and do not show it; a consumer parsing the CSV with `csv` sees it,
# which is why the count goes in the sidecar rather than staying silent.
CSV_FORMULA_ESCAPE = "'"


def _csv_escape(v: Any) -> Any:
"""Neutralise a formula-shaped cell. Returns (value, was_escaped).

Only STRINGS are touched. Escaping a number would turn `-5` into the text
`'-5` and break every sum a consumer writes over the column.

CSV only. The JSON output keeps the site's bytes exactly as served, so
the two files deliberately differ; `csv_cells_escaped` in the sidecar is
what declares that divergence instead of leaving it to be discovered.
"""
if isinstance(v, str) and v.startswith(CSV_FORMULA_LEADS):
return CSV_FORMULA_ESCAPE + v, True
return v, False


def write_csv(rows: Sequence[Any], path: str, row_cls: Type = Profile) -> int:
# An empty result still gets the header row. A zero-byte file makes a
# consumer fail on read (no columns to parse) instead of reading a valid
# table with zero rows — and "an empty result is still a well-formed
Expand All @@ -370,11 +435,21 @@ def write_csv(rows: Sequence[Any], path: str, row_cls: Type = Profile) -> None:
# The header comes from `row_cls`, not from the first row, so an empty
# run still writes the columns of the mode that produced it.
fieldnames = [f.name for f in fields(row_cls)]
escaped = 0
with _atomic(path, newline="") as f:
writer = csv.DictWriter(f, fieldnames=fieldnames)
writer.writeheader()
for r in rows:
writer.writerow({k: _csv_value(v) for k, v in asdict(r).items()})
row = {}
for k, v in asdict(r).items():
# Escape AFTER `_csv_value`: a list joined into one cell is
# text too, and its first element can be formula-shaped
# while the list itself is not a string.
value, was_escaped = _csv_escape(_csv_value(v))
row[k] = value
escaped += was_escaped
writer.writerow(row)
return escaped


# Exit code used when a run completes but produced nothing. Distinct from 1
Expand Down Expand Up @@ -515,7 +590,8 @@ def run_meta(status: str, stop_reason: str, pages_requested: int,


def save(rows: Sequence[Any], out_prefix: str, fmt: str,
allow_empty: bool = False, row_cls: Type = Profile) -> int:
allow_empty: bool = False, row_cls: Type = Profile,
stats: Optional[dict] = None) -> int:
"""Write JSON/CSV and return a process exit code.

Returns 0 when rows were written, EXIT_NO_PRODUCTS when there were none.
Expand All @@ -542,8 +618,17 @@ def save(rows: Sequence[Any], out_prefix: str, fmt: str,
write_json(rows, f"{out_prefix}.json")
print(f"[+] Saved {len(rows)} rows -> {out_prefix}.json")
if fmt in ("csv", "both"):
write_csv(rows, f"{out_prefix}.csv", row_cls=row_cls)
escaped = write_csv(rows, f"{out_prefix}.csv", row_cls=row_cls)
print(f"[+] Saved {len(rows)} rows -> {out_prefix}.csv")
# Reported through `stats` rather than as a return value because
# this function's return IS the exit code.
if stats is not None:
stats["csv_cells_escaped"] = escaped
if escaped:
print(f"[!] {escaped} CSV cell(s) began with a formula character "
f"and were prefixed with {CSV_FORMULA_ESCAPE!r} so a "
f"spreadsheet reads them as text. The JSON output is "
f"unchanged — see csv_cells_escaped in the sidecar.")
return 0 if rows else EXIT_NO_PRODUCTS


Expand Down Expand Up @@ -604,8 +689,14 @@ def finish_run(rows: Sequence[Any], out_prefix: str, fmt: str,
# failed, the run is not complete, whatever it stopped for.
complete = stop_reason in COMPLETE_STOP_REASONS and not pages_failed
row_cls = ROW_CLASS_BY_MODE.get(mode, Profile)
rc = save(rows, out_prefix, fmt, allow_empty=allow_empty, row_cls=row_cls)
stats: dict = {}
rc = save(rows, out_prefix, fmt, allow_empty=allow_empty, row_cls=row_cls,
stats=stats)
wrote_output = bool(rows) or allow_empty
# Housekeeping goes UNDER the caller's extra, never over it: a name
# collision must not let a count of ours silently replace a fact the
# engine recorded about the site.
extra = {**stats, **(extra or {})}

if wrote_output:
status = "complete" if (rows and complete) else (
Expand Down
40 changes: 39 additions & 1 deletion playwright_scraper.py
Original file line number Diff line number Diff line change
Expand Up @@ -1029,6 +1029,7 @@ def _run_profile(session_box, pw, args, pool) -> Tuple[List[Any], Dict[str, Any]
more_beyond_page: List[str] = []
empty_slots = 0
countries = set()
drift: Dict[str, List[str]] = {}
blocked = False
for i, outcome in enumerate(outcomes):
blocked = blocked or outcome.blocked
Expand All @@ -1046,6 +1047,8 @@ def _run_profile(session_box, pw, args, pool) -> Tuple[List[Any], Dict[str, Any]
empty_slots += int(diag.get("spotlight_empty_slots") or 0)
if diag.get("viewer_country"):
countries.add(diag["viewer_country"])
for key in diag.get("payload_keys_missing") or ():
drift.setdefault(key, []).append(handles[i])
# `page` is WHICH ACCOUNT of the run a row came from, so that
# `page`+`position` is unique across a multi-account run — CLAUDE.md
# §18 records a sibling where 60 of 119 rows claimed a position
Expand Down Expand Up @@ -1088,6 +1091,12 @@ def _run_profile(session_box, pw, args, pool) -> Tuple[List[Any], Dict[str, Any]
# how a reader checks that a proxy or a Scraping Browser
# `country-` segment did what it was asked.
"viewer_countries": sorted(countries) or None,
# Keys this parser reads that a served page no longer carried, with
# the accounts each was missing on. Empty on every page this repo
# was built and tested against; non-empty means Snapchat moved a
# field and some columns of those accounts' rows are nulls of the
# parser's making, not the account's. The canary fails on it.
"payload_keys_missing": drift or None,
}
return rows, meta

Expand All @@ -1111,7 +1120,23 @@ def __enter__(self):
def __exit__(self, *exc):
return self._pw.__exit__(*exc)

def _transport_used(args) -> str:
"""What actually fetched the pages: "http", "browser" or "cdp".

`--transport` says what was ASKED for, and the default `auto` is not a
transport at all — the sidecar used to record "auto" for a run that
never started a browser, so a reader could not tell from it whether
Chromium had been involved (third-party audit, 2026-10-05). `auto`
switches `args.transport` to "browser" the moment it falls back, so
reading it after the run gives the answer.
"""
if args.cdp_endpoint:
return "cdp"
return "browser" if getattr(args, "transport", "auto") == "browser" else "http"


def scrape(args) -> int:
transport_requested = getattr(args, "transport", "auto")
pool = proxy_pool_from_args(args)
handles = _handles(args)
if pool and args.concurrency > 1:
Expand Down Expand Up @@ -1141,7 +1166,12 @@ def scrape(args) -> int:
"blocked")}
extra["engine"] = "playwright"
extra["category"] = args.category
extra["transport"] = getattr(args, "transport", "auto")
# `engine` names the CLI that ran (which driver the browser path would
# use); `transport` names what actually fetched the pages. The audit
# asked for `engine=http` instead — but the engine is still this
# script, and folding the two into one field would lose which one.
extra["transport"] = _transport_used(args)
extra["transport_requested"] = transport_requested

unavailable = meta.get("handles_unavailable")
if unavailable:
Expand All @@ -1153,6 +1183,14 @@ def scrape(args) -> int:
logger.info("%d account(s) have no public profile, so Snapchat shows "
"them no Spotlight and no story: %s", len(not_public),
", ".join("@" + h for h in not_public))
drift = meta.get("payload_keys_missing")
if drift:
logger.warning("Snapchat's page no longer carries %d key(s) this "
"parser reads: %s. The page was served and parsed, "
"but columns read from those keys are null for the "
"parser's reasons, not the account's. See "
"payload_keys_missing in the sidecar.", len(drift),
", ".join(sorted(drift)))
more = meta.get("handles_with_more_than_page")
if more and args.mode != "profile":
logger.info("%d account(s) have more Spotlight/highlights than their "
Expand Down
Loading
Loading