Predict Motor Octane Number of Hydrocarbon Mixtures
This tutorial shows how to predict the Motor Octane Number (MON) of a hydrocarbon mixture. The dataset stems from the paper from Chew et al.. It contains 722 experiments of up to 121 componet mixtures of 423 individual molecules. Each molecule is represented by its SMILES string.
/opt/hostedtoolcache/Python/3.12.14/x64/lib/python3.12/site-packages/torch/jit/_script.py:1491: FutureWarning: `torch.jit.script` is deprecated. Please switch to `torch.compile` or `torch.export`.
warnings.warn(
/opt/hostedtoolcache/Python/3.12.14/x64/lib/python3.12/site-packages/torch/jit/_script.py:1491: FutureWarning: `torch.jit.script` is deprecated. Please switch to `torch.compile` or `torch.export`.
warnings.warn(
Setup Data
df_experiments = get_octane_data()df_experiments["valid_MON"] =1output_key ="MON"inputs = Inputs( features = [ ContinuousInput(key=col, bounds=(0, 1), descriptors=Descriptors(structure=[col]))for col in df_experiments.columnsif col notin ["MON", "Label", "valid_MON"] ])outputs = Outputs(features=[ContinuousOutput(key=output_key)])
/tmp/ipykernel_3093/2040360654.py:2: PerformanceWarning: DataFrame is highly fragmented. This is usually the result of calling `frame.insert` many times, which has poor performance. Consider joining all columns at once using pd.concat(axis=1) instead. To get a de-fragmented frame, use `newframe = frame.copy()`
df_experiments["valid_MON"] = 1
Setup Surrogate and perform CV
We model the high-dimensional problem by using an engineered feature called WeightedSumFeature. It computes the weighted sum of the molecular descriptors of the original ContinuousInputs (each carrying a structure) that make up the engineered feature. Here we use Mordred descriptors with a correlation cutoff of 0.9.
surrogate_data = SingleTaskGPSurrogate( inputs=inputs, outputs=outputs, engineered_features = EngineeredFeatures( features=[ WeightedSumFeature( key="mixture", features=inputs.get_keys(), columns=[], generators=[ MordredDescriptors(descriptors=mordred_names, ignore_3D=False) ], filter_descriptors=True, correlation_cutoff=0.9, keep_features=False, ) ] ))print("Number of molecular features before correlation filtering: ", len(surrogate_data.engineered_features[0].generators[0].get_descriptor_names()))surrogate = surrogates.map(surrogate_data)cv_train, cv_test, _ = surrogate.cross_validate(df_experiments, folds=10ifnot SMOKE_TEST else3)display(cv_test.get_metrics())print("Number of molecular features before correlation filtering: ", len(surrogate_data.engineered_features[0].generators[0].get_descriptor_names()))
Number of molecular features before correlation filtering: 1826
/opt/hostedtoolcache/Python/3.12.14/x64/lib/python3.12/site-packages/bofire/utils/torch_tools.py:51: UserWarning: The given NumPy array is not writable, and PyTorch does not support non-writable tensors. This means writing to this tensor will result in undefined behavior. You may want to copy the array to protect its data or make it writable before converting it to a tensor. This type of warning will be suppressed for the rest of this program. (Triggered internally at /__w/pytorch/pytorch/torch/csrc/utils/tensor_numpy.cpp:213.)
return torch.from_numpy(np.ascontiguousarray(df.to_numpy())).to(**tkwargs)
MAE
MSD
R2
MAPE
PEARSON
SPEARMAN
FISHER
0
3.623217
47.914292
0.849359
3.175749e+14
0.92181
0.917017
6.251366e-123
Number of molecular features before correlation filtering: 1826
Even better performance can be achieved by using SAAS based surrogates, like AdditiveMapSaasSingleTaskGPSurrogate or EnsembleMapSaasSingleTaskGPSurrogate. Drawback are higher computational costs.
---title: Predict Motor Octane Number of Hydrocarbon Mixturesjupyter: python3---This tutorial shows how to predict the Motor Octane Number (MON) of a hydrocarbon mixture. The dataset stems from the paper from [Chew et al.](https://www.nature.com/articles/s41524-025-01552-2). It contains 722 experiments of up to 121 componet mixtures of 423 individual molecules. Each molecule is represented by its SMILES string.## Imports```{python}import osimport bofire.surrogates.api as surrogatesfrom bofire.benchmarks.data.octane_number import get_octane_datafrom bofire.data_models.domain.api import EngineeredFeatures, Inputs, Outputsfrom bofire.data_models.features.api import ( ContinuousInput, ContinuousOutput, Descriptors, WeightedSumFeature,)from bofire.data_models.descriptor_generators.api import MordredDescriptorsfrom bofire.data_models.descriptor_generators.names import mordred as mordred_namesfrom bofire.data_models.surrogates.api import SingleTaskGPSurrogateSMOKE_TEST = os.environ.get("SMOKE_TEST")```## Setup Data```{python}df_experiments = get_octane_data()df_experiments["valid_MON"] =1output_key ="MON"inputs = Inputs( features = [ ContinuousInput(key=col, bounds=(0, 1), descriptors=Descriptors(structure=[col]))for col in df_experiments.columnsif col notin ["MON", "Label", "valid_MON"] ])outputs = Outputs(features=[ContinuousOutput(key=output_key)])```## Setup Surrogate and perform CVWe model the high-dimensional problem by using an engineered feature called `WeightedSumFeature`. It computes the weighted sum of the molecular descriptors of the original `ContinuousInput`s (each carrying a `structure`) that make up the engineered feature. Here we use Mordred descriptors with a correlation cutoff of 0.9.```{python}surrogate_data = SingleTaskGPSurrogate( inputs=inputs, outputs=outputs, engineered_features = EngineeredFeatures( features=[ WeightedSumFeature( key="mixture", features=inputs.get_keys(), columns=[], generators=[ MordredDescriptors(descriptors=mordred_names, ignore_3D=False) ], filter_descriptors=True, correlation_cutoff=0.9, keep_features=False, ) ] ))print("Number of molecular features before correlation filtering: ", len(surrogate_data.engineered_features[0].generators[0].get_descriptor_names()))surrogate = surrogates.map(surrogate_data)cv_train, cv_test, _ = surrogate.cross_validate(df_experiments, folds=10ifnot SMOKE_TEST else3)display(cv_test.get_metrics())print("Number of molecular features before correlation filtering: ", len(surrogate_data.engineered_features[0].generators[0].get_descriptor_names()))```Even better performance can be achieved by using SAAS based surrogates, like `AdditiveMapSaasSingleTaskGPSurrogate` or `EnsembleMapSaasSingleTaskGPSurrogate`. Drawback are higher computational costs.