Skip to content

About

No description, website, or topics provided.

Resources

Stars

4 stars

Watchers

1 watching

Forks

Latest commit

 

History

13 Commits

Folders and files

Repository files navigation

scraper-591

Overview

A rental-listing scraper for 591租屋網. Fetches a filtered search page with Requests and parses its listing cards with BeautifulSoup.

The scraper exports the first results page as UTF-8 JSON. It validates listing IDs, titles, links, and monthly rents, and removes duplicate listing IDs. See Extraction verification for findings and verification scope.

本工具用於擷取 591 租屋資料,使用 Requests 取得指定篩選條件的搜尋頁面,再透過 BeautifulSoup 解析房源卡片。

程式將搜尋結果第一頁匯出為 UTF-8 JSON,驗證房源 ID、標題、連結及月租金,並排除重複的房源 ID。查核結果與範圍另見資料擷取查核紀錄(英文)。

Setup and usage

Install uv and use Python 3.13 or later. From the repository root:

uv sync --frozen
uv run main.py

The default search requests Taipei whole-home listings with 3 or 4+ rooms, pets allowed, and monthly rent of NT$30,000–40,000. These are requested filters, not a guarantee that every returned listing matches them.

To use another search or output file:

uv run main.py --url "https://rent.591.com.tw/list?region=1&kind=1" --output data/custom.json

Use uv run main.py --help for available options.

Output

The default output is data\listing.json. Parent directories are created automatically. A successful run replaces the output file; an HTTP or parsing failure exits with status 1 without replacing existing output.

Each array entry contains:

Field Contents
id, title, url Listing ID (string), title, and link
price, price_unit Rent as displayed, such as "38,000" and "元/月"
extra_fee_text Displayed fee note, or an empty string if absent
details Displayed information rows, including layout, area, floor, location, and publisher information when present
tags Listing tags, such as 可養寵物

Prices and descriptive text retain their source formatting. Fee notes are not converted into a total monthly cost, and detail rows are not split into normalized fields.

Tests and checks

Pytest is included in the development dependencies installed by uv sync --frozen.

uv run pytest
uv run ruff check .
uv run ruff format --check .

Unit tests use synthetic HTML and mocked HTTP responses. End-to-end tests run the real CLI against a local HTTP server, covering successful output and failure handling. Neither contacts 591. To run an optional live end-to-end check, use a separate output path:

uv run main.py --output data/live-check.json

A live run depends on the site's availability and current markup; it is not part of the offline test suite.

Limitations and troubleshooting

  • Only one results page is fetched. There is no automatic pagination or detail-page extraction.
  • Browser-like request headers worked during verification but do not guarantee access. HTTP errors and timeouts are reported; there is no automatic retry or browser fallback.
  • Missing cards, missing required fields, unreadable prices, or unexpected rent units produce an error rather than silently exporting incomplete results. An actual zero-result search is not yet distinguished from changed markup or a blocked response.
  • The parser depends on the site's HTML structure. If extraction fails, inspect the response and update the parser and tests.
  • Search-filter correctness and optional descriptive fields are not fully validated. There is no scheduling, cross-run deduplication, or listing history.

Contributing and responsible use

Bug reports, feature requests, and pull requests are welcome. Include reproduction steps and relevant errors without posting private data. Add offline regression tests for parser changes.

Review the site's terms and applicable requirements before collecting data. Keep request volume modest, respect access restrictions, and avoid unnecessary collection of personal information.

About

No description, website, or topics provided.

Resources

Stars

4 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages