Supervisor

Dr. Muhammad Iqbal

Programme

BSc (Hons) in Computing in IT

Subject

Computer Science

Abstract

Deforestation has become a major environmental issue worldwide, especially in the Amazon Rainforest, which is known as one of the world’s largest carbon sinks. This region has been significantly impacted by cattle farming, contributing to increased CO₂ emissions and land degradation. As Brazil is one of the world’s largest beef producers and exporters, the environmental impact of livestock production has attracted increasing attention from organisations, governments, and sustainability analysts.

This project aims to analyse the relationship between cattle farming, deforestation, and CO₂ emissions in the Amazon using Machine Learning. Using the CRISP-DM framework, the project explored each stage of the process, starting with business understanding, followed by data understanding, data preparation, modelling, evaluation, and deployment.

During the exploratory data analysis phase, the dataset was analysed using histograms, scatter plots, boxplots, and correlation analysis to better understand the distribution of the variables, identify outliers, examine relationships between features, and detect issues such as skewness and multicollinearity. The analysis showed that the dataset contains highly skewed distributions, extreme values, correlated variables, and missing data, which required careful preprocessing before modelling.

Numerous data preparation techniques were applied, including missing value handling, removal of invalid entries, label encoding, and removal of unnecessary columns. Three different approaches for handling missing values were tested: dropping missing values, median imputation, and model-based imputation.

Four Machine Learning models were developed and compared: K-Nearest Neighbours Regressor (KNN), Random Forest Regressor, Partial Least Squares Regressor (PLS) and Linear Regression. Feature scaling and PCA were applied to some models. Different hyperparameters were tested using GridSearchCV, and the models were evaluated using metrics such as R² score, Mean Absolute Error (MAE), and cross-validation. Particular attention was given to how correlated features and preprocessing techniques affected model performance and generalisation.

Date of Award

2026

Full Publication Date

2026

Access Rights

open access

Document Type

Undergraduate Project

Resource Type

bachelor thesis

Share

COinS