CORTEXA
← Browse
openalexZenodo (CERN European Organization for Nuclear Research)Cited by 0

Financial data source preparation and analysis as the initial stage of stochastic time series modeling by machine learning techniques

Yurchenko Yuriy, Oleksandr Zakovorotnyi

Proceedings of the Scientific Conference "The 13th International Scientific and Practical Online Conference of Young Scientists and Students ‘Contemporary Problems of Automation and Control’".The conference talk presents a structured approach to preparing and analyzing financial data sources as a prerequisite for applying machine learning methods to stochastic modeling of financial time series. The focus is on forming a research environment that connects real market data providers, open ML datasets, and modern software tools such as MATLAB and Python for subsequent forecasting and risk modeling tasks.First, the report reviews key categories of financial data sources, including API-based market data providers (such as Financial Modeling Prep and Yahoo!Finance) and open machine learning platforms (Kaggle, Hugging Face) that supply benchmark datasets for credit risk and market analysis. It compares their structure, data availability, and suitability for stochastic modeling tasks, emphasizing issues of data quality, feature-target definition, and preprocessing requirements for time series and tabular financial data.Next, the talk demonstrates practical examples of building machine learning pipelines on this data: credit default prediction on the HMEQ dataset in MATLAB using Decision Trees, and Apple stock price modeling in Python using Yahoo!Finance data and Random Forest regressors. The performance of these models is evaluated with standard metrics (ROC, AUC, Accuracy, Precision, MAE, MAPE, MSE), showing how properly prepared data and carefully selected sources directly influence the accuracy and robustness of stochastic time series models in financial applications.

Also available via: European Organization for Nuclear Research

View free PDFSource page

Related papers

openalexZenodo (CERN European Organization for Nuclear Research)2026-07-26

Enhancing Cardiovascular Disease Diagnosis through Data-Driven Feature Analysis and Cross-Validated Machine Learning Models

Abhilash Butola

Abstract - Cardiovascular diseases are a major global health problem, accounting for 17.9 million deaths per year and constituting 32 percent globally. According to the World Health Organization, the disease in people is due to an unhealthy diet,such as the intake of more junk fo…

View free PDFSource page
openalexZenodo (CERN European Organization for Nuclear Research)2026-07-23

A Scalable Distributed and Fault-Tolerant Architecture for Cloud-Based Machine Learning and Data Analysis

Grace Dooshima GBOR, Emmanuel Ogala, Donald Douglas Atsa’am, Iorshashe Agaji

Abstract The rapid growth of data-intensive applications has necessitated the development of scalable and efficient architectures for cloud-based machine learning and data analysis. This study proposes a scalable, distributed, and fault-tolerant architecture designed to address t…

View free PDFSource page
openalexZenodo (CERN European Organization for Nuclear Research)2026-07-26

Comparative Analysis of Machine Learning Classification Algorithms and Hybrid Models for Student Performance Prediction

Ms. Pooja C. Soni, Dr. Hetal R. Modi, PC Negi

This study focuses on the analysis and comparison of machine learning classification algorithms and hybrid machine learning models for predicting student academic performance. Educational Data Mining techniques are used to extract meaningful insights from student datasets. Variou…

View free PDFSource page
openalexZenodo (CERN European Organization for Nuclear Research)2026-07-24

Before the Model: Why Datasets and Data Representation Define What Machine Learning Can Learn

Jean Franck Loa Rojas

Machine learning systems do not learn reality directly; they learn from the representations preserved in their datasets. This structured narrative review examines how dataset purpose, coverage, integrity, labeling, independence, reproducibility, governance, and continuity determi…

View free PDFSource page
openalexZenodo (CERN European Organization for Nuclear Research)2026-07-24

Simple Data Cleaning Techniques for Improving Machine Learning

Jyoti Panthangi

High-quality data is essential for building reliable machine learning models. Raw datasets often contain missing values, outliers, duplicates, inconsistent formats, and unstructured categorical variables. These issues reduce model accuracy and lead to biased predictions. This pap…

View free PDFSource page
openalexZenodo (CERN European Organization for Nuclear Research)2026-07-23

Stop Spatializing Time: Machine Learning Agents Should Learn Through Time, Not About Time

Teeratham Vitchutripop, Alyssa Quarles, Wei Zhang, Daniel Rakita

Modern machine learning systems are increasingly deployed in settings that require persistent interaction, adaptation, memory, and decision-making over time. Yet, most learning paradigms remove the temporal pressures faced by physically embedded agents: the world waits for comput…

View free PDFSource page