{"repo":"amitkedia007/Financial-Fraud-Detection-Using-LLMs","free":true,"listed":false,"github":"https://github.com/amitkedia007/Financial-Fraud-Detection-Using-LLMs","clone":"git clone https://github.com/amitkedia007/Financial-Fraud-Detection-Using-LLMs.git","description":"The aim of this dissertation is to assess the effectiveness of LLMs such as FinBERT and GPT-2 in detecting fraudulent activities in financial reports and statements. This repo provides the code for implementing LLMs, traditional machine learning and deep learning models on the labelled dataset","language":"Jupyter Notebook","stars":88,"topics":["classification","deep-learning","finance","financial-fraud-detection","large-language-models","machine-learning"],"license":null,"category":"machine-learning","readme_excerpt":"Financial Fraud Detection Using AI Get Deeper Understanding from motivation to data collection to model training to final results you can checkout the project blogs here: https://www.amitkedia.com/project/67ce1a818013ee818192b171 Project Overview Introduction: This project utilizes machine learning, deep learning, and Large Language Models (LLMs) to detect financial fraud. It's based on a comprehensive dataset derived from financial filings to the U.S. Securities and Exchange Commission (SEC), aiming to compare and enhance AI models in identifying fraudulent financial activities. Objective: The goal is to foster a collaborative platform where data scientists and researchers can develop, test, and improve AI models for detecting financial fraud. Dataset Description Source: The dataset includes financial filings from 170 companies, split equally between those involved in fraudulent and non-fraudulent activities. Structure: Each dataset entry contains details such as Central Index Key (CIK), filing year, company name, and a categorical indicator of fraud. Final Dataset: Finally the dataset is out on Kaggle do check it out here.. Data Preprocessing Preprocessing steps involve text cleaning, tokenization, and transforming data into machine-readable formats, ensuring balanced and fair model training. Model Implementation The project encompasses a variety of models, including Logistic Regression, SVM, Random Forest, XGBoost, ANN, HAN, GPT-2, and FinBERT, selected for their NLP capab","default_branch":null,"files":null,"tree":[],"storefront":"/r/amitkedia007","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/amitkedia007/Financial-Fraud-Detection-Using-LLMs/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}