You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
This repository was archived by the owner on Sep 8, 2026. It is now read-only.
Updates to readme documentation
Updates to application in crawl.py
Selenium module upgrade
Compatibility fix for running this Django 4.0 app with Python 13+ (created cgi.py)
Upgraded Web Crawling Pipeline (CI)
Copy file name to clipboardExpand all lines: readme.md
+26-19Lines changed: 26 additions & 19 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -6,8 +6,14 @@
6
6
7
7
*[Important note:](#note)
8
8
9
+
*[Requirements](#requirements)
10
+
9
11
*[Installation](#installation)
10
12
13
+
*[Configure This Project](#configure-this-project)
14
+
15
+
*[Run Your Selenium Server Jar File](#run-your-selenium-server-jar-file)
16
+
11
17
*[Usage](#usage)
12
18
13
19
*[Using Docker](#using-docker)
@@ -30,54 +36,55 @@
30
36
31
37
## Important note: <aname="note"></a>
32
38
33
-
Before you try to scrape any website, go through its robots.txt file. You can access it via `domainname/robots.txt`. There, you will see a list of pages allowed and disallowed for scraping. You should not violate any terms of service of any website you scrape.
39
+
Before you try to scrape any website, go through its robots.txt file. You can access it via `domainname/robots.txt`. There, you will see a list of pages allowed and disallowed for scraping. You should not violate any terms of service of any website you scrape.
40
+
41
+
## Requirements
42
+
43
+
*[Tested using Python 3.13](https://www.python.org)
[CLI options in the Selenium Grid](https://www.selenium.dev/documentation/grid/configuration/cli_options/).
69
+
60
70
## Usage
61
71
62
72
[XPath cheat sheet](https://devhints.io/xpath).
63
73
64
-
Update the command at [crawl.py](https://github.com/kkamara/python-selenium/blob/main/seleniumpy/management/commands/crawl.py)
74
+
Update the command at [crawl.py](https://github.com/kkamara/python-selenium/blob/main/seleniumpy/management/commands/crawl.py) to perform your instructions in web scraping.
65
75
66
76
```bash
67
-
alias py="python3"
68
-
py manage.py crawl
77
+
python manage.py crawl
69
78
```
70
79
71
-
If you still need help installing and running the app check out the readme at https://github.com/kkamara/python-react-boilerplate which is the base system for this python-selenium app.
72
-
73
80
## Using Docker?
74
81
75
82
```bash
76
83
alias compose='docker-compose -f local.yml'
77
84
compose build
78
85
compose up
79
86
# Automated runs with Docker:
80
-
# compose up --build -d && python3 manage.py crawl
0 commit comments