A Flask-based machine-learning web application developed as an Ideonix Internship Project. It accepts a loan dataset, trains a classifier to predict loan eligibility, and provides a form for making individual predictions.
- Project Overview
- Problem Statement and Objectives
- Features
- Technology Stack
- Machine Learning Workflow
- Architecture
- Application Workflow
- Project Structure
- Dataset
- Model and Evaluation
- Requirements and Setup
- Running the Application
- Using the Application
- Results and Visual Artifact
- Limitations and Future Enhancements
- Learning Outcomes
- Author
- License
The application brings together a small Flask web interface and a scikit-learn training and inference flow. A user can register and log in, upload a CSV file, preview its first rows, train a model, and enter applicant details to receive an eligibility classification.
The project stores account records in a MySQL database. Uploaded datasets, the serialized model, and the generated confusion-matrix image are stored as local files.
Loan eligibility decisions depend on several applicant and loan attributes. This project demonstrates how those attributes can be used in a supervised classification workflow and exposed through a browser-based interface.
Project objectives:
- Accept a loan dataset in CSV format and preview its contents.
- Train a classifier using selected applicant and loan attributes.
- Report held-out accuracy and generate a confusion matrix after training.
- Accept one applicant record through a web form and display a predicted class.
- Provide basic account registration and login backed by MySQL.
- Registration and login, with submitted passwords hashed using Flask-Bcrypt.
- CSV upload and a preview of up to 10 rows.
- Dataset validation for the target and feature columns required by the training code.
- Random Forest training and serialization to
models/model.pkl. - Accuracy reporting and confusion-matrix image generation during training.
- A single-record eligibility form and results page.
- Server-rendered pages styled with a local stylesheet and Bootstrap CSS loaded from a CDN.
| Area | Technologies present in the project |
|---|---|
| Language and web framework | Python, Flask, Jinja templates |
| Data and machine learning | pandas, NumPy, scikit-learn, joblib |
| Evaluation and visualization | scikit-learn metrics, Matplotlib, Seaborn |
| Persistence and password hashing | Flask-SQLAlchemy, MySQL through PyMySQL, Flask-Bcrypt |
| Front end | HTML, CSS, Bootstrap 5.0.2 stylesheet via CDN |
The repository does not include a dependency manifest or a pinned Python version. The install commands below are derived from the imports and database driver used in app.py.
- Upload a CSV file. The upload route saves it under
uploads/and previews its first 10 rows. - Training checks for
Loan_Statusand selects the eight features listed below. Loan_Statusis mapped to1when its string value isY(case-insensitive, with surrounding whitespace removed); other values are mapped to0.- Object-valued selected features are transformed with scikit-learn's
LabelEncoder. Feature values are converted to numeric values, with conversion failures replaced by0. - The data is split into training and test subsets using an 80/20 split and
random_state=42. - A
RandomForestClassifierwith 100 estimators andrandom_state=42is fit on the training subset. - The model is saved to
models/model.pkl. The test subset is used to calculate accuracy and create a confusion matrix. - The prediction route loads that saved model and predicts a class from the eight form values.
flowchart LR
A[CSV upload] --> B[Preview and schema check]
B --> C[Select features and encode]
C --> D[80/20 train-test split]
D --> E[Fit Random Forest]
E --> F[Save models/model.pkl]
E --> G[Calculate accuracy and confusion matrix]
H[Applicant form] --> I[Load saved model]
F --> I
I --> J[Eligibility result]
flowchart TB
Browser[Browser]
Flask[Flask routes and Jinja templates]
Uploads[(Local uploads directory)]
Model[(Serialized model file)]
Plot[(Confusion-matrix image)]
DB[(MySQL user table)]
Browser <--> Flask
Flask --> Uploads
Flask --> Model
Flask --> Plot
Flask <--> DB
The Flask routes coordinate both the web pages and the ML operations. User records use the configured SQLAlchemy database; uploaded CSVs, the trained model, and the plot are file-based. The training route uses the uploaded dataframe held in the running Flask process, so upload and training are expected to happen in the same app process.
- Open the home page and register an account, or sign in with an existing account.
- From the dashboard, choose the loan-eligibility workflow.
- Upload a CSV file and review the displayed preview.
- Select Train Model to train and save the classifier.
- Select Detect, enter the requested applicant details, and submit the form.
- Review the predicted eligibility label and the submitted values.
The dashboard route checks for a logged-in session. The upload, training, and detection routes do not currently enforce that session check.
.
├── app.py # Flask routes, database model, upload, training, and prediction logic
├── Loan_Eligibility.csv # Included dataset (500 data rows)
├── models/
│ └── model.pkl # Serialized model artifact
├── static/
│ ├── css/
│ │ └── style.css # Application styles
│ └── images/
│ ├── confusion_matrix.png # Existing training-plot artifact
│ ├── images.jpg # Referenced as the page background
│ ├── img.jpg # Image asset; purpose not established by app.py
│ └── My First Blog.html # Standalone HTML file, not used by the Flask routes
├── templates/
│ ├── base.html # Shared page layout and flash messages
│ ├── index.html # Home page
│ ├── login.html # Login form
│ ├── register.html # Registration form
│ ├── dashboard.html # Signed-in dashboard
│ ├── upload.html # CSV upload and preview
│ ├── train.html # Training result view
│ ├── detect.html # Applicant input form
│ └── result.html # Prediction result
└── uploads/
└── Loan_Eligibility.csv # Upload-directory copy of the included dataset
The root and uploads/ CSV files are identical in the current workspace. Files under uploads/ can be replaced when a CSV with the same filename is uploaded.
The repository includes Loan_Eligibility.csv, which contains 500 data rows and 12 columns. The same dataset is also present under uploads/. Its origin and collection methodology are not documented in the repository.
| Column | Role in the current training code |
|---|---|
Gender |
Present in dataset; not selected for training |
Married |
Present in dataset; not selected for training |
Dependents |
Selected feature |
Education |
Present in dataset; not selected for training |
Employment_Type |
Selected feature |
ApplicantIncome |
Selected feature |
CoapplicantIncome |
Selected feature |
LoanAmount |
Selected feature |
Loan_Term_Months |
Selected feature |
Credit_History |
Selected feature |
Property_Area |
Selected feature |
Loan_Status |
Target (Y becomes 1; other values become 0) |
Training expects all eight selected feature columns and Loan_Status. A CSV that only has a different or partial schema will not satisfy the current training code.
The training route creates a scikit-learn RandomForestClassifier with n_estimators=100 and random_state=42. It evaluates the fitted model on the 20% test split using accuracy and a confusion matrix. No cross-validation or additional metrics are implemented in the current training route.
The prediction form submits these features: ApplicantIncome, CoapplicantIncome, LoanAmount, Loan_Term_Months, Credit_History, Employment_Type, Dependents, and Property_Area.
Encoding caveat: training applies LabelEncoder to object-valued features, but the prediction form submits manually assigned integer option values. The form's Employment_Type codes are not in the alphabetical order that LabelEncoder uses for the dataset's text labels. The app does not persist or reuse fitted encoders, so this category can be encoded inconsistently between training and prediction. Treat predictions accordingly until the encoding is made consistent.
The checked-in models/model.pkl is a serialized artifact, but the repository does not record which dataset or training run produced it. Running the training workflow replaces it.
- Python installed and available as
python(the project does not specify a supported Python version). - MySQL server available locally.
- A MySQL database named
ml_appand connection settings configured for the local environment. - Internet access when loading the Bootstrap stylesheet from its CDN.
The application currently defines its database connection directly in app.py; update that configuration for your local MySQL instance before running. Do not commit real credentials or secret keys. The repository does not currently provide environment-variable configuration or a .env file.
From the project root, in Windows PowerShell:
py -m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install --upgrade pipInstall the packages imported by the application:
python -m pip install Flask Flask-SQLAlchemy Flask-Bcrypt pandas numpy scikit-learn joblib matplotlib seaborn PyMySQLCreate the MySQL database before starting the application:
CREATE DATABASE ml_app;On startup, app.py calls db.create_all() to create the configured user table. Dependency versions are not pinned in this repository.
Start the Flask development server from the project root:
python app.pyThen open http://127.0.0.1:5000. The application is configured to listen on port 5000 and all host interfaces. Running from the project root is important because model and plot paths are relative to the working directory.
- Register or log in to reach the dashboard.
- Upload a
.csvfile. The upload handler accepts the CSV extension and displays at most the first 10 records. - Train using a dataset with the required columns described above.
- After training, use the detection form to submit the eight feature values and view the predicted label.
The form provides choices for credit history, employment type, dependents, and property area, and numeric inputs for income, loan amount, and term. Its numeric limits and display labels are defined in templates/detect.html.
The training route reports an accuracy percentage and writes a confusion-matrix plot when training completes. The exact score depends on the uploaded dataset and the current split; this README does not claim a fixed result. No verified benchmark report or application UI screenshots are included in the repository. The existing image below is the confusion-matrix artifact, not a product screenshot; it may reflect a previous training run.
- Use a fitted preprocessing pipeline so categorical encoders are persisted and shared by training and prediction.
- Move database settings and Flask's secret key to environment-based configuration.
- Add consistent authentication and authorization checks to upload, training, and prediction routes.
- Add upload-size limits, stronger CSV/schema validation, and clearer input error handling.
- Add dependency/version management and automated tests for routes, preprocessing, and prediction.
- Record dataset provenance and repeatable evaluation results before publishing model-quality claims.
- Add genuine screenshots of the running application if portfolio screenshots are desired.
These are suggested improvements, not features currently implemented by the project.
The code demonstrates practical work with Flask routes and Jinja templates, form handling and CSV uploads, SQLAlchemy-backed user records, password hashing, pandas preprocessing, scikit-learn model training and evaluation, joblib serialization, and Matplotlib/Seaborn visualization.
Ideonix Internship Project
The author's name and profile are not identified in the repository. Add the preferred name and portfolio or GitHub link here before publishing.
No license file or license declaration is present in the repository. The project's reuse and distribution terms need to be confirmed by the author.
