Scientific Products
ML Explorer
ML Explorer by Atomistica
Machine Learning Made Accessible for Scientific Research and Education
ML Explorer is a browser-based machine learning platform developed by Atomistica to support scientific research, education, and data-driven modeling in chemistry, physics, materials science, and related disciplines.
Designed for researchers, students, educators, and professionals, ML Explorer provides an intuitive environment for preparing datasets, training machine learning models, evaluating their performance, interpreting predictions, and applying trained models to new data.
The platform combines essential machine learning methods with advanced analytical capabilities, allowing users to perform comprehensive machine learning workflows without writing code or installing specialized software.
Data Preparation and Preprocessing
Prepare scientific datasets for machine learning using integrated data processing tools:
Dataset Cleaning – Identify and remove duplicate records, missing values, and invalid entries.
Exploratory Data Analysis (EDA) – Examine dataset characteristics, descriptive statistics, distributions, and correlations.
Outlier Detection – Identify and handle potential outliers using Z-score and interquartile range (IQR) methods.
Categorical Encoding – Convert categorical variables into numerical representations using label encoding and one-hot encoding.
Data Transformation – Apply logarithmic, square-root, Box–Cox, and Yeo–Johnson transformations to numerical variables.
CSV Utilities – Prepare and standardize datasets for further analysis.
Machine Learning for Regression
Develop predictive models for continuous numerical properties using established machine learning algorithms:
Linear Regression – Model linear relationships between input features and target variables.
Ridge Regression – Apply L2 regularization to linear regression models.
Lasso Regression – Apply L1 regularization for model regularization and feature selection.
Decision Tree Regression – Model nonlinear relationships through decision-based learning.
Random Forest Regression – Combine multiple decision trees for ensemble-based predictions.
Support Vector Regression (SVR) – Model complex relationships using support vector methods and kernel functions.
The regression module supports customizable model parameters, optional feature scaling, training and testing datasets, and performance evaluation using R², MAE, MSE, and RMSE.
Machine Learning for Classification
Build models for predicting categorical outcomes using widely applied classification algorithms:
Logistic Regression – Perform classification using probabilistic linear models.
k-Nearest Neighbors (kNN) – Classify samples based on neighboring observations.
Support Vector Classification (SVC) – Classify data using support vector methods and kernel functions.
Decision Tree Classification – Develop interpretable decision-based classification models.
Random Forest Classification – Apply ensemble learning for robust classification tasks.
Users can configure model parameters, select input features and target variables, and evaluate classification performance on training and testing datasets.
Model Validation and Performance Evaluation
Assess model performance and generalization using integrated evaluation tools:
Train/Test Splitting – Divide datasets into independent training and testing subsets.
Cross-Validation – Evaluate regression and classification models using K-Fold, Stratified K-Fold, and repeated cross-validation approaches, as applicable.
Performance Metrics – Evaluate predictive accuracy and model performance using task-appropriate statistical metrics.
Graphical Evaluation – Examine model performance through visualizations and diagnostic plots.
Optional Feature Scaling – Apply standardized feature scaling within supported machine learning workflows.
These capabilities help users compare models, assess predictive reliability, and identify potential overfitting.
Model Interpretation and Feature Importance
Understand how machine learning models make predictions using integrated interpretability tools:
Permutation Feature Importance – Evaluate the contribution of individual input features by measuring changes in model performance.
SHAP Analysis – Interpret predictions using SHapley Additive exPlanations.
SHAP Summary Visualizations – Explore feature contributions across datasets.
Feature-Level Interpretation – Identify variables that most strongly influence model predictions.
Classification-Specific SHAP Analysis – Examine feature contributions for individual target classes.
These tools support more transparent and scientifically meaningful interpretation of machine learning results.
Model Export and Predictions
Extend machine learning workflows beyond initial model training:
Model Export – Save trained models together with the preprocessing information required for subsequent predictions.
Prediction on New Data – Apply trained models to new samples using compatible input features.
Interactive Predictions – Enter feature values directly to obtain predictions.
Inverse Target Transformations – Convert supported transformed regression predictions back to their original physical scale.
Model Reuse – Reuse exported models within supported prediction workflows.
These capabilities are particularly useful for predicting molecular properties, material characteristics, experimental outcomes, and other scientific quantities.
Why ML Explorer?
No programming required – Build and evaluate machine learning models through an intuitive graphical interface.
No installation required – Access machine learning tools directly through a web browser.
Complete workflows – From raw datasets and preprocessing to model training, interpretation, and prediction.
Scientifically relevant methods – Established regression, classification, validation, and interpretability techniques.
Transparent results – Detailed performance metrics and visual explanations support informed scientific interpretation.
Flexible applications – Suitable for chemistry, physics, materials science, and other data-driven research fields.
Research and education – Designed for practical scientific investigations, classroom demonstrations, and independent learning.
Continuous development – Regular improvements and additional capabilities based on scientific and educational needs.
ML Explorer by Atomistica — Making Machine Learning Accessible for Scientific Discovery.
