Skip to content

Releases: TelecomsXChangeAPi/OpenTextShield

v2.11.0 — obfuscation and SMPP bypass fixes (model 2.7)

Choose a tag to compare

@ajamous ajamous released this 18 Sep 21:52
196f6c6

OpenTextShield v2.11.0 — Obfuscation and SMPP bypass fixes (model 2.7 unchanged)

Released. Platform version 2.11.0, promoted from v2.11.0-rc.2 after local validation and a soak test. rc.2 superseded rc.1:
rc.1 carried an incorrect claim about the Docker build context and lacked the ignore-rule
and Git LFS fixes below. The shipped
classifier stays model 2.7.

Headline

Attackers could get past OpenTextShield without beating the model at all: by
writing in fullwidth letters, by adding a header the proxy trusted, by padding a
message past the API's length limit, or by hiding the text in a second field.
This release closes those paths and stops the proxy from delivering unsure
verdicts as safe. No model change, so no regression risk on classification.

What changed

1. Disguised text is cleaned properly (enhanced_preprocessing.py)

Obfuscation of the same phishing messages (n=544) Before After
Fullwidth letters (PayPal) 42.6% blocked 100%
Look-alike letters outside the old map (ӏ ѕ ԁ) 80.9% blocked 100%
All styles combined 92.1% 99.3%

The old folding also damaged real text: it rewrote Russian and Greek messages
into mixed-script nonsense and mapped Cyrillic у to u through a duplicate
key. Folding is now context-aware, and on ~19,000 benchmark messages the new
cleanup changes zero predictions, where the old one flipped 18.

2. The SMPP proxy no longer delivers unsure verdicts as safe

Below confidence_threshold the proxy forwarded every spam or phishing verdict
as ham. Across all saved benchmark predictions that delivered 564 IMC25 phishing
messages and prevented only 5 false blocks. The new below_threshold_action
defaults to use_label; set it to forward_as_ham for the old behaviour.

3. Classification can no longer be skipped by a sender

  • A single-shift header on text with no escape characters, and plain-ASCII
    ISO-2022-JP payloads, are now classified instead of skipped.
  • Shift headers on non-GSM data coding are flagged rather than trusted.
  • Messages over 512 characters are fitted to the API limit (start and end kept)
    instead of failing the call and being forwarded unclassified.
  • short_message and message_payload are classified together when both are
    present, so a harmless short message cannot hide a phishing payload.
  • New unclassifiable_action lets strict operators reject what cannot be read.
  • Config is validated at startup, and malformed API replies no longer leave a
    client without a submit_sm_resp.

4. Evaluation now matches production

run_eval.py and calibrate_thresholds.py tokenised raw text while the API
normalised first, so published numbers did not describe the deployed pipeline
for non-ASCII traffic. Both now apply the same cleanup (--raw-text reproduces
older numbers). Training does too, via finetune_tier1.py --train-csv.

5. Honest eval sets

  • fable5_adversarial_v1_clean.csv (71 rows): the suite without rows that share
    a template with the synthetic training set. The previous token-Jaccard check
    missed obfuscated and translated twins; two rows had digit-for-digit copies in
    training.
  • hard_legit_a2p_v1.csv (40 rows): legitimate bank, delivery, billing and code
    messages that look like scams, to measure false blocks. Handwritten synthetic.

6. Training data cleaned and documented (not shipped)

docs/LABELING_GUIDE.md plus evals/label_audit.py produce a relabeled,
deduplicated training set. Retraining on it gives a real trade rather than a
free win, so model 2.7 stays. Full evidence, including three seeds per
configuration, is in evals/results/DATA_CLEANUP_2.8.md.

7. Docker images ship only what they should

Docker's ignore patterns do not cross /, so a bare *.pth only ever matched
root-level files. Model 2.7 was therefore never excluded, and an earlier claim in
this branch that images built from main ship without the model was wrong: a real
docker build disproved it. The same gap did ship three things it should not:

In the image before Size Why it matters
mbert_ots_model_2.8-candidate.pth 711 MB an unreleased, unvalidated checkpoint
audit_logs/ varies full SMS text (audit_text_storage defaults to full)
infra/ 1.3 GB Terraform state and provider binaries

Patterns now cross directories (**/*.pth, **/node_modules/, **/__pycache__/)
with explicit re-includes for the shipped 2.7 and the 2.5 rollback target, and
audit_logs/, feedback/ and infra/ are excluded. Verified against the real
engine: build context 2.8 GB to 1.4 GB, image 4.25 GB containing exactly 2.5 and 2.7,
no audit logs, no infra.

8. The loader refuses a Git LFS pointer

This is the real way an image can start and be unable to classify. A checkout
without git lfs pull leaves a ~130 byte pointer file where the weights belong; it
passes the existence check, so the service used to start and fail later inside
torch.load. The loader now detects the pointer signature, logs what to run, and
skips the model rather than pretending to serve it.

Tests

Suite Result
API unit tests 39/39
SMPP offline (npm test) 137/137
SMPP integration against model 2.7 47/47 and 32/32, identical to the previous proxy

Soak test on 2026-09-18

18,482 labeled messages through the proxy in 1.2 hours (about 4 messages/s with
hourly bursts): 0 errors or timeouts, container memory flat at 1.5 GB, proxy RSS
flat at 30 MB, latency p50 145 ms, p99 under 470 ms, health checks all green.
Spam blocked 97.7%, phishing 85.1%, fullwidth obfuscation 14/14. False blocks:
2.0% on personal ham, 12.5% on synthetic A2P notices (the known model 2.7
weakness). Full report: SOAK_v2.11.0-rc.2.md.

Validated locally on 2026-09-18

Run on an Apple Silicon machine against the branch tip, model 2.7:

Claim Result
API reports platform 2.11.0, model 2.7 confirmed, host and container
Fullwidth lure blocked phishing 1.000 (same verdict as the plain-ASCII text)
Real Russian and Greek undamaged ham 0.999 / 0.982; normaliser output inspected directly
Unsure verdicts acted on 10 below-threshold spam/phishing verdicts kept their label and were rejected
Long message no longer skipped 661 characters via message_payload: classified phishing 0.9999, rejected
Hidden second field no longer skipped harmless short_message + phishing message_payload: rejected
Startup config validation unknown rule action and unknown below_threshold_action both exit 1 with a clear error
Integration suites 47/47 and 32/32
Unit suites API 42/42, SMPP offline 137/137
fable5 clean block rate 55/55 = 100%, false blocks 1/16
hard legit blocked 13/40
Mishra block 98.8%, false blocks 0.5% (26/4844)
Docker image builds, serves model 2.7, blocks the fullwidth lure

Two probes returned ham where a reader might expect otherwise: "Log in to
раураӏ now" and a math-bold PayPal lure. The normaliser folds both correctly to
ASCII; model 2.7 simply does not flag those short texts, and the plain-ASCII
versions score identically. That is a model limitation, not a regression, and it is
what the data work in DATA_CLEANUP_2.8.md targets.

Upgrade notes

  • Behaviour change: unsure spam and phishing verdicts are now acted on. If
    your traffic makes false blocks the bigger risk, set
    classification.below_threshold_action to forward_as_ham in config.json.
  • Startup is now strict: an unknown rule action, previously a silent hang,
    stops the proxy with a clear error.
  • No model file changes, so no Git LFS pull is required for this release.

Validating

No images are published for a release candidate. Build locally from the branch;
Docker Hub tags follow only after validation and merge.

git checkout data-cleanup-step3
git lfs pull                                    # the .pth must be a real file
cd src/smpp_interface && npm test               # 137 offline tests
cd ../.. && python -m pytest src/api_interface/tests/ -q

# local image, and confirm it serves the model it claims
docker compose up --build -d
curl -s localhost:8002/health                   # api_version 2.11.0, model 2.7
curl -s -X POST localhost:8002/predict/ -H 'Content-Type: application/json' \
  -d '{"text":"Your parcel is held, pay the fee: parcel-fee.top","model":"ots-mbert"}'

# benchmarks against the honest eval sets
python evals/run_eval.py \
  --model src/mBERT/training/model-training/mbert_ots_model_2.7.pth \
  --dataset fable5:evals/datasets/fable5_adversarial_v1_clean.csv \
  --dataset csv:evals/datasets/hard_legit_a2p_v1.csv --tag v2.11.0-rc1

On live traffic, the number to watch is the false-block rate on A2P alerts:
unsure spam and phishing verdicts are now acted on instead of delivered.

Docker Hub

Multi-arch image (linux/amd64, linux/arm64) built from this tag:

docker pull telecomsxchange/opentextshield:2.11.0   # digest sha256:5dceb88dba67c306f75f83e7d31010260e0a4676935f09cb5e431be08e691737
docker pull telecomsxchange/opentextshield:latest   # same image

Both architectures were pulled from Hub and verified to report platform 2.11.0 with model 2.7 and to block a fullwidth-obfuscated lure.

v2.11.0-rc.2 — validated locally (supersedes rc.1)

Choose a tag to compare

@ajamous ajamous released this 18 Sep 19:31

OpenTextShield v2.11.0 — Obfuscation and SMPP bypass fixes (model 2.7 unchanged)

Release candidate (v2.11.0-rc.2). Platform version 2.11.0. rc.2 supersedes rc.1:
rc.1 carried an incorrect claim about the Docker build context and lacked the ignore-rule
and Git LFS fixes below. The shipped
classifier stays model 2.7. This RC is for validation on its own branch and
is not a production release.

Headline

Attackers could get past OpenTextShield without beating the model at all: by
writing in fullwidth letters, by adding a header the proxy trusted, by padding a
message past the API's length limit, or by hiding the text in a second field.
This release closes those paths and stops the proxy from delivering unsure
verdicts as safe. No model change, so no regression risk on classification.

What changed

1. Disguised text is cleaned properly (enhanced_preprocessing.py)

Obfuscation of the same phishing messages (n=544) Before After
Fullwidth letters (PayPal) 42.6% blocked 100%
Look-alike letters outside the old map (ӏ ѕ ԁ) 80.9% blocked 100%
All styles combined 92.1% 99.3%

The old folding also damaged real text: it rewrote Russian and Greek messages
into mixed-script nonsense and mapped Cyrillic у to u through a duplicate
key. Folding is now context-aware, and on ~19,000 benchmark messages the new
cleanup changes zero predictions, where the old one flipped 18.

2. The SMPP proxy no longer delivers unsure verdicts as safe

Below confidence_threshold the proxy forwarded every spam or phishing verdict
as ham. Across all saved benchmark predictions that delivered 564 IMC25 phishing
messages and prevented only 5 false blocks. The new below_threshold_action
defaults to use_label; set it to forward_as_ham for the old behaviour.

3. Classification can no longer be skipped by a sender

  • A single-shift header on text with no escape characters, and plain-ASCII
    ISO-2022-JP payloads, are now classified instead of skipped.
  • Shift headers on non-GSM data coding are flagged rather than trusted.
  • Messages over 512 characters are fitted to the API limit (start and end kept)
    instead of failing the call and being forwarded unclassified.
  • short_message and message_payload are classified together when both are
    present, so a harmless short message cannot hide a phishing payload.
  • New unclassifiable_action lets strict operators reject what cannot be read.
  • Config is validated at startup, and malformed API replies no longer leave a
    client without a submit_sm_resp.

4. Evaluation now matches production

run_eval.py and calibrate_thresholds.py tokenised raw text while the API
normalised first, so published numbers did not describe the deployed pipeline
for non-ASCII traffic. Both now apply the same cleanup (--raw-text reproduces
older numbers). Training does too, via finetune_tier1.py --train-csv.

5. Honest eval sets

  • fable5_adversarial_v1_clean.csv (71 rows): the suite without rows that share
    a template with the synthetic training set. The previous token-Jaccard check
    missed obfuscated and translated twins; two rows had digit-for-digit copies in
    training.
  • hard_legit_a2p_v1.csv (40 rows): legitimate bank, delivery, billing and code
    messages that look like scams, to measure false blocks. Handwritten synthetic.

6. Training data cleaned and documented (not shipped)

docs/LABELING_GUIDE.md plus evals/label_audit.py produce a relabeled,
deduplicated training set. Retraining on it gives a real trade rather than a
free win, so model 2.7 stays. Full evidence, including three seeds per
configuration, is in evals/results/DATA_CLEANUP_2.8.md.

7. Docker images ship only what they should

Docker's ignore patterns do not cross /, so a bare *.pth only ever matched
root-level files. Model 2.7 was therefore never excluded, and an earlier claim in
this branch that images built from main ship without the model was wrong: a real
docker build disproved it. The same gap did ship three things it should not:

In the image before Size Why it matters
mbert_ots_model_2.8-candidate.pth 711 MB an unreleased, unvalidated checkpoint
audit_logs/ varies full SMS text (audit_text_storage defaults to full)
infra/ 1.3 GB Terraform state and provider binaries

Patterns now cross directories (**/*.pth, **/node_modules/, **/__pycache__/)
with explicit re-includes for the shipped 2.7 and the 2.5 rollback target, and
audit_logs/, feedback/ and infra/ are excluded. Verified against the real
engine: build context 2.8 GB to 1.4 GB, image 4.25 GB containing exactly 2.5 and 2.7,
no audit logs, no infra.

8. The loader refuses a Git LFS pointer

This is the real way an image can start and be unable to classify. A checkout
without git lfs pull leaves a ~130 byte pointer file where the weights belong; it
passes the existence check, so the service used to start and fail later inside
torch.load. The loader now detects the pointer signature, logs what to run, and
skips the model rather than pretending to serve it.

Tests

Suite Result
API unit tests 39/39
SMPP offline (npm test) 137/137
SMPP integration against model 2.7 47/47 and 32/32, identical to the previous proxy

Validated locally on 2026-09-18

Run on an Apple Silicon machine against the branch tip, model 2.7:

Claim Result
API reports platform 2.11.0, model 2.7 confirmed, host and container
Fullwidth lure blocked phishing 1.000 (same verdict as the plain-ASCII text)
Real Russian and Greek undamaged ham 0.999 / 0.982; normaliser output inspected directly
Unsure verdicts acted on 10 below-threshold spam/phishing verdicts kept their label and were rejected
Long message no longer skipped 661 characters via message_payload: classified phishing 0.9999, rejected
Hidden second field no longer skipped harmless short_message + phishing message_payload: rejected
Startup config validation unknown rule action and unknown below_threshold_action both exit 1 with a clear error
Integration suites 47/47 and 32/32
Unit suites API 42/42, SMPP offline 137/137
fable5 clean block rate 55/55 = 100%, false blocks 1/16
hard legit blocked 13/40
Mishra block 98.8%, false blocks 0.5% (26/4844)
Docker image builds, serves model 2.7, blocks the fullwidth lure

Two probes returned ham where a reader might expect otherwise: "Log in to
раураӏ now" and a math-bold PayPal lure. The normaliser folds both correctly to
ASCII; model 2.7 simply does not flag those short texts, and the plain-ASCII
versions score identically. That is a model limitation, not a regression, and it is
what the data work in DATA_CLEANUP_2.8.md targets.

Upgrade notes

  • Behaviour change: unsure spam and phishing verdicts are now acted on. If
    your traffic makes false blocks the bigger risk, set
    classification.below_threshold_action to forward_as_ham in config.json.
  • Startup is now strict: an unknown rule action, previously a silent hang,
    stops the proxy with a clear error.
  • No model file changes, so no Git LFS pull is required for this release.

Validating this RC

No images are published for a release candidate. Build locally from the branch;
Docker Hub tags follow only after validation and merge.

git checkout data-cleanup-step3
git lfs pull                                    # the .pth must be a real file
cd src/smpp_interface && npm test               # 137 offline tests
cd ../.. && python -m pytest src/api_interface/tests/ -q

# local image, and confirm it serves the model it claims
docker compose up --build -d
curl -s localhost:8002/health                   # api_version 2.11.0, model 2.7
curl -s -X POST localhost:8002/predict/ -H 'Content-Type: application/json' \
  -d '{"text":"Your parcel is held, pay the fee: parcel-fee.top","model":"ots-mbert"}'

# benchmarks against the honest eval sets
python evals/run_eval.py \
  --model src/mBERT/training/model-training/mbert_ots_model_2.7.pth \
  --dataset fable5:evals/datasets/fable5_adversarial_v1_clean.csv \
  --dataset csv:evals/datasets/hard_legit_a2p_v1.csv --tag v2.11.0-rc1

On live traffic, the number to watch is the false-block rate on A2P alerts:
unsure spam and phishing verdicts are now acted on instead of delivered.

v2.11.0-rc.1 — obfuscation and SMPP bypass fixes (model 2.7)

Choose a tag to compare

@ajamous ajamous released this 18 Sep 06:46

Superseded by v2.11.0-rc.2, which corrects an incorrect claim about the Docker build context and adds the ignore-rule and Git LFS pointer fixes found during local validation. Do not validate against rc.1.

OpenTextShield v2.10.0-rc.1 - Smarter Smishing Detection (model 2.7)

Choose a tag to compare

@ajamous ajamous released this 14 Jun 04:35

OpenTextShield v2.10.0 — Smarter Smishing Detection (model 2.7)

Release candidate (v2.10.0-rc.1). Platform version 2.10.0, shipping the
new mBERT classifier model 2.7. This RC is for validation on its own branch
and is not a production release.

Headline

Model 2.7 is a validated, overfitting-aware fine-tune continued from v2.5. It
beats v2.5 on the block rate of all four benchmarks with no classic-spam
regression
, and dramatically improves phishing recall — especially on modern,
real-world, and adversarial smishing where the previous model was leaking attacks
to users.

Benchmark results — v2.5 → model 2.7

Benchmark n 3-class acc Block rate Phishing recall
UCI SMS Spam (classic, public) 5,574 99.3% → 99.5% 99.3% → 99.5% n/a (no phishing class)
Mishra & Soni Phishing (public) 5,971 88.8% → 89.4% 99.2% → 99.3% 2.7% → 6.1%
IMC25 smishing (real-world, public) 8,007 20.4% → 47.0% 69.9% → 72.4% 17.1% → 45.7%
Adversarial suite (in-house) 150 48.7% → 86.7% 73.3% → 96.7% 11.5% → 80.8%
  • Block rate up on every benchmark — including UCI, which ticks up rather
    than regressing, confirming the classic-spam behaviour is fully protected.
  • Phishing recall transformed on the real-world and adversarial sets: IMC25
    17.1% → 45.7%, adversarial 11.5% → 80.8%. The previously silent
    phishing → ham leaks (family impersonation, vishing callback, toll/delivery-fee
    scams, and text obfuscation) are now caught.

Validation note: IMC25 was used to select the training hyperparameters, so treat
its 72.4% as a validation result. UCI, Mishra & Soni, and the adversarial suite
are independent confirmations — all improved.

The recipe

Continue from the v2.5 base over a targeted synthetic dataset plus a 2,500/class
rehearsal sample of the original corpus; plain cross-entropy loss, lr 1e-5,
2 epochs
. Heavier configurations (class-weighted/focal loss, more epochs, larger
lr) scored higher on the in-domain validation split but overfit and regressed the
real-world IMC25 block rate; the gentle recipe recovered generalization.

Other changes in this release

  • Tier-1 training pipeline: validation split with best-checkpoint selection
    (saves the best epoch by validation macro-F1, not the last), and a per-run
    training log capturing the full metric curve.
  • Threshold-calibration utility (evals/calibrate_thresholds.py) to tune the
    phishing/spam decision boundary via a per-class logit bias without retraining.
  • Obfuscation hardening in the live path: homoglyph and zero-width-character
    normalization is now applied in the dynamic-batching inference path, not just the
    per-request preprocessor — closing obfuscation evasion independent of the model.
  • API wired to model 2.7 and api_version set to 2.10.0.
  • Reproducible evaluation harness (evals/run_eval.py, compare_runs.py)
    with the v2.5 and v2.7 summary evidence committed under evals/results/.

Artifact

  • mbert_ots_model_2.7.pth — bert-base-multilingual-cased, 3-class
    (ham / spam / phishing), 104+ languages.

How to test

git checkout sms-classification-intelligence-eval
git lfs pull
source ots/bin/activate
python evals/run_eval.py \
  --model src/mBERT/training/model-training/mbert_ots_model_2.7.pth \
  --dataset uci:/tmp/uci.tsv \
  --dataset mishra:evals/datasets/mishra_soni_5971.csv \
  --dataset imc25:/tmp/imc25.csv:8000 --dataset fable5 --tag v2.7
python evals/compare_runs.py \
  --before evals/results/summary_v2.5.json --after evals/results/summary_v2.7.json

OpenTextShield v2.5.0 - Enhanced Multilingual AI Security Platform

Choose a tag to compare

@ajamous ajamous released this 06 Oct 02:46
d9736a6

🚀 OpenTextShield v2.5.0 Release

✨ Key Features

  • Enhanced Multilingual Support: Comprehensive testing across 10 languages (English, Spanish, French, German, Chinese, Arabic, Hebrew, Russian, Japanese, Hindi) with 99.9% accuracy
  • PEFT/LoRA Adapters: Parameter-efficient fine-tuning for Hebrew language support and incremental learning
  • TMForum API Compliance: Full TMF922 AI Inference Job Management API for enterprise integrations
  • Improved Security: Multi-stage Docker builds, non-root execution, distroless options
  • Incremental Learning: Continuous model improvement with new training data

🐳 Docker Images

    • Standard production image
    • Enhanced security with multi-stage builds
    • Latest stable release

📊 Performance Improvements

  • Superior multilingual classification accuracy vs v2.1.0
  • Optimized Apple Silicon MLX support
  • Reduced container sizes with secure builds (6.75GB vs 13GB)

🔧 Technical Enhancements

  • Modular API architecture with FastAPI
  • Comprehensive error handling and logging
  • Async job processing for TMForum endpoints
  • Professional web interface with real-time monitoring

🧪 Testing

  • Comprehensive multilingual comparison testing
  • Stress testing up to 20k requests
  • API integration tests
  • Docker deployment verification

📚 Documentation

  • Updated CLAUDE.md with development commands
  • Enhanced README with deployment options
  • API documentation and OpenAPI specs

🤝 Enterprise Features

  • TMForum-compliant API endpoints
  • Async inference job management
  • Professional frontend interface
  • Production-ready security posture

For deployment instructions, see README.md and CLAUDE.md.

v2.1.1

Choose a tag to compare

@ajamous ajamous released this 28 Jun 08:35


🚀 OpenTextShield v2.1.1 – Modernization & Security Release

🎯 Highlights

  • Complete platform modernization

  • Major security enhancements

  • Streamlined deployment and modular architecture

  • No breaking changes — full backward compatibility


🏗️ Platform Architecture Overhaul

✅ Modular API Redesign (src/api_interface/)

  • Enterprise-grade, clean structure:

    • config/ – Env-based configs

    • middleware/ – IP verification, CORS

    • models/ – Pydantic schemas with validation

    • routers/ – Endpoints (health, prediction, feedback)

    • services/ – Business logic

    • utils/ – Logging & exceptions

  • Dependency injection, versioned APIs, structured logging, and robust error handling

✅ Codebase Cleanup

  • Removed:

    • Legacy BERT/FastText/mBERT models

    • 4.7GB+ of archived models and training sets

    • Redundant requirements files


🔐 Security Enhancements

🔧 Dependency Updates (Critical & High)

  • h11 → v0.16.0 (chunked encoding)

  • Starlette → v0.46.2 (DoS)

  • Transformers → v4.53.0 (deserialization fix)

  • protobuf → v5.28.3 (DoS)

  • ecdsa → v0.19.0 (timing attack)

  • urllib3, setuptools, localstack – patched

🐳 Docker Security Levels

Dockerfile | Security | Highlights -- | -- | -- Dockerfile.secure | 🛡️ Recommended | Multi-stage, non-root, smaller image (~2.27GB) Dockerfile.distroless | 🔒 Max Security | No shell/pkg manager, ultra-minimal Dockerfile | 📦 Enhanced Standard | Patched base, system updates

🙏 Contributors & Roadmap

🔍 Thanks to:

  • Security reviewers

  • Codebase modernizers

  • Test engineers

  • Documentation contributors

🛣️ What’s Next:

  • Model speed/accuracy tuning

  • Continuous security upgrades

  • Deployment integrations

  • API enhancements

Read more

OTS 2.1

Choose a tag to compare

@ajamous ajamous released this 09 May 10:14

OTS 2.1.0

Open Text Shield (OTS) Version 2.1 - Release Notes

🌟 Exciting News! We're thrilled to announce OTS Version 2.1! 🌟

This update marks a significant enhancement as we extend our language support to include Sinhala and Tamil, further broadening the accessibility and versatility of the OTS models.

🚀 Highlights

🇱🇰 Sinhala Data Set: Expanding our linguistic reach to better serve the Sinhala-speaking communities.
🇮🇳 Tamil Data Set: Strengthening our capabilities in Tamil, enriching user interactions and data understanding.

🌍 Language Support Expansion

With Version 2.1, OTS now proudly offers text classification predictions across 10 diverse languages, reaffirming our commitment to breaking language barriers and enhancing global communication:

🇬🇧 English
🇸🇦 Arabic
🇮🇩 Indonesian
🇩🇪 German
🇮🇹 Italian
🇪🇸 Spanish
🇷🇺 Russian
🇫🇷 French
🇱🇰 Sinhala
🇮🇳 Tamil

As always, we are committed to continuous improvement and are excited to see how our latest enhancements will empower our users. Thank you for your continued support and trust in OTS!

OTS 2.0.0

Choose a tag to compare

@ajamous ajamous released this 31 Mar 14:48

Open Text Shield (OTS) Version 2.0 - Release Notes

🌟 We're excited to unveil OTS Version 2.0! 🌟

This latest update broadens our horizon with the integration of three additional languages, enriching the OTS experience and expanding our global footprint.


🚀 Highlights

  • 🇷🇺 Russian Data Set: Elevating our understanding of the Russian language.
  • 🇪🇸 Spanish Data Set: Diving deeper into the richness of Spanish.
  • 🇫🇷 French Data Set: Embracing the elegance of French.

🌍 Language Support Expansion

With Version 2.0, OTS now proudly offers text classification predictions across 8 diverse languages:

  1. 🇬🇧 English
  2. 🇸🇦 Arabic
  3. 🇮🇩 Indonesian
  4. 🇩🇪 German
  5. 🇮🇹 Italian
  6. 🇪🇸 Spanish
  7. 🇷🇺 Russian
  8. 🇫🇷 French

.

Full Changelog: 1.9.0...2.0.0

OTS 1.9.0 - Support for German/Italian Languages

Choose a tag to compare

@ajamous ajamous released this 28 Mar 12:26

Release Notes for OTS Version 1.9.0

We are thrilled to announce the release of OTS Version 1.9.0, which is trained on two additional languages.

Highlights

  • Trained on Italian data set
  • Trained on German data set

OTS 1.9.0 can predict text classification for 5 languages (English, Arabic, Indonesian, German, Italian)

What's Changed

Full Changelog: 1_7_5...1.9.0

OTS 1.7.5 - Support for Indonesian Language

Choose a tag to compare

@ajamous ajamous released this 19 Mar 14:46

Release Notes for OTS Version 1.7.5

We are thrilled to announce the release of OTS Version 1.7.5, which brings new language support.

Highlights

  • Full Indonesian Dataset Integration: We have fully integrated a comprehensive Indonesian dataset to expand our coverage and improve processing, understanding, and insights for ID text. This enhancement aims to cater to our diverse global audience, offering better support for Indonesian language content.

Full Changelog: 1_7...1_7_5