Monorepo prototype for the Scicommons web app.
Current deployment shape:
React frontend -> FastAPI backend -> Postgres user database
|
v
Article/search service -> SQLite + FAISS article artifacts
For the prototype, all services can run on one Arbutus server.
backend/ FastAPI user/session/profile API and article-service proxy
frontend/ React/Vite frontend
scicomm_embedding/ Article fetching, indexing, semantic/keyword search API
igather2/ PubMed ingestion support and bundled EDirect tools
docker-compose.yml Local Postgres user database
setup.sh Install dependencies and initialize local configuration
run.sh Run Postgres, article service, backend, and frontend
The article database is intentionally separate from the user database. User
state lives in Postgres. Article records and search indexes live in
scicomm_embedding artifacts such as all.sqlite, all_specter.index,
all_metadata.json, and all_manifest.json.
SSH to the Arbutus server at 134.87.8.193.
Clone the repo:
git clone git@github.com:m2b3/test_integration.git
cd test_integrationIf GitHub SSH is not set up on the server, add a server SSH key to GitHub first.
Run setup from the repo root:
./setup.shOn the Arbutus server, set the browser-facing backend URL before setup so the
frontend .env points at the reachable backend:
VITE_API_BASE_URL=http://134.87.8.193:8000 ./setup.shThen run everything:
./run.shOpen:
http://134.87.8.193:5173
Local defaults:
Frontend: http://localhost:5173
Backend: http://localhost:8000
Article service: http://localhost:8100
Postgres: localhost:5432
setup.sh does the following:
- creates local
.envfiles for the root scripts, backend, frontend,scicomm_embedding, andigather2 - creates Python virtual environments for
backend,scicomm_embedding, andigather2 - installs backend, article pipeline/search, and PubMed ingestion dependencies
- runs
npm installinfrontend - starts Postgres through Docker Compose
- recreates and seeds the user database
- installs a user crontab entry that runs the article pipeline nightly at 2 AM
Useful environment overrides:
VITE_API_BASE_URL=http://134.87.8.193:8000 ./setup.sh
NCBI_EMAIL=you@example.com NCBI_API_KEY=optional_key ./setup.sh
OVERWRITE_ENV=1 VITE_API_BASE_URL=http://134.87.8.193:8000 ./setup.sh
OVERWRITE_FRONTEND_ENV=1 VITE_API_BASE_URL=http://134.87.8.193:8000 ./setup.sh
RESET_USER_DB=0 ./setup.sh
RUN_PIPELINE=1 ./setup.sh
INSTALL_ARTICLE_CRON=0 ./setup.sh
ARTICLE_CRON_SCHEDULE="0 2 * * *" ./setup.shRESET_USER_DB=1 is the default and runs backend/setup_database.py, which
drops and recreates the prototype user-side tables. RUN_PIPELINE=0 is the
default because the article pipeline performs network fetching and embedding
work.
The generated .env files are intentionally still ignored by Git. Use
OVERWRITE_ENV=1 ./setup.sh to regenerate all of them from the current
environment variables.
By default, setup installs CPU-only PyTorch before the article search
dependencies. This avoids pulling large CUDA wheels on ordinary CPU servers.
If a previous setup failed with No space left on device during torch
installation, clear the partial article environment and pip cache before
rerunning:
rm -rf scicomm_embedding/.venv
python3 -m pip cache purge || true
OVERWRITE_ENV=1 ./setup.shRun the article pipeline when you need fresh article artifacts:
cd scicomm_embedding
source .venv/bin/activate
python pipeline.pyThe article service expects these artifacts in scicomm_embedding/ by default:
all.sqlite
all_specter.index
all_metadata.json
all_manifest.json
If artifacts live elsewhere, point run.sh at them:
SCICOMM_ARTIFACT_DIR=/path/to/artifacts ./run.shrun.sh starts:
- Postgres via Docker Compose
- article/search API on port
8100 - backend API on port
8000 - frontend dev server on port
5173
run.sh loads the root .env before applying defaults, so changes to ports,
database URLs, artifact paths, or service URLs can be made there after setup.
It also frees the configured service ports before starting the services and
prints any process IDs it stops. Set KILL_PORTS=0 to disable this behavior.
Useful environment overrides:
START_DB=0 ./run.sh
KILL_PORTS=0 ./run.sh
BACKEND_PORT=8001 ARTICLE_PORT=8101 FRONTEND_PORT=5174 ./run.sh
VITE_API_BASE_URL=http://134.87.8.193:8000 ./run.shVerify from the server:
curl http://localhost:8000/health
curl http://localhost:8000/tags
curl http://localhost:8100/healthVerify externally:
curl http://134.87.8.193:8000/healthIf local curl works but external curl hangs, open inbound TCP ports 8000 and
5173 in the Arbutus security/firewall settings.
The scripts default to sudo docker compose, matching the Arbutus server setup.
If your local Docker user does not need sudo, use:
DOCKER_COMPOSE="docker compose" ./setup.sh
DOCKER_COMPOSE="docker compose" ./run.shor add the user to the Docker group and reconnect:
sudo usermod -aG docker $USERsetup.sh installs this user crontab block by default:
# scicommons nightly article pipeline start
0 2 * * * cd /path/to/test_integration && ./scripts/nightly_article_pipeline.sh
# scicommons nightly article pipeline end
The cron job runs scripts/nightly_article_pipeline.sh, which loads the
generated env files, activates scicomm_embedding/.venv, and runs
python pipeline.py. Logs go to:
scicomm_embedding/logs/nightly_pipeline.log
To inspect or remove the cron entry:
crontab -l
crontab -eBackend:
cd backend
source .venv/bin/activate
DATABASE_URL=postgresql://scicommons:scicommons@localhost:5432/scicommons \
ARTICLE_SERVICE_BASE_URL=http://localhost:8100 \
uvicorn app.main:app --host 0.0.0.0 --port 8000Article service:
cd scicomm_embedding
source .venv/bin/activate
uvicorn article_service.main:app --host 0.0.0.0 --port 8100Frontend:
cd frontend
npm run dev -- --host 0.0.0.0 --port 5173Use tmux if you want services to keep running after logout:
tmux new -s scicommons
./run.shDetach from tmux:
Ctrl+b, then d
Reconnect later:
tmux attach -t scicommonsUseful tmux commands:
tmux ls
tmux kill-session -t scicommonsGET /health
GET /tags
GET /sources
GET /articles
GET /articles?semantic_query=biology&source=arxiv
GET /users/{user_id}/feed
GET /users/{user_id}/profile
PUT /users/{user_id}/profile
GET /users/{user_id}/tags
GET /users/{user_id}/recently-viewed
POST /users/{user_id}/recently-viewed
POST /login
GET /me
POST /logout
PUT /users/{user_id}/tags
scicomm_embedding and igather2 were imported with non-squashed Git subtree
merges, so their histories are reachable from this repository's history.
For future work in the old standalone repositories, do not force-push over a partner's branch. Have collaborators commit or stash local work before pulling updates:
git status
git stash push -u -m "work before monorepo sync"
git pull
git stash popCommitted work in the old repositories can still be brought into this monorepo
with git subtree pull for the appropriate prefix.
- Backend runs on port
8000. - Article/search service runs on port
8100by default. - Frontend dev server runs on port
5173. - Postgres runs on local port
5432. - User login/session/profile/interests/recently viewed are stored through the backend and Postgres. The browser also keeps a non-authoritative cache of the last profile and source filter so refreshes feel remembered;
/meverifies the session and refreshes profile data from the DB. - Article listing/search is proxied through
ARTICLE_SERVICE_BASE_URL; Postgres is not the source of article records. - Frontend, user backend, Postgres, and article/search service can run together for now; they can be split across servers once deployment requirements are clearer.