data_models.kernels.api

data_models.kernels.api

Classes

Name Description
Kernel Covariance function of a Gaussian process.
AdditiveKernel Sum of several kernels, \(k(\mathbf x, \mathbf x') = \sum_i k_i(\mathbf x, \mathbf x')\).
MultiplicativeKernel Product of several kernels, \(k(\mathbf x, \mathbf x') = \prod_i k_i(\mathbf x, \mathbf x')\).
ScaleKernel Wraps another kernel with a fitted output scale,
MaternKernel Matern kernel, a less smooth alternative to the RBF kernel.
RBFKernel Radial basis function kernel, the usual default for continuous inputs.
LinearKernel Linear kernel, equivalent to Bayesian linear regression on the inputs.
PolynomialKernel Polynomial kernel, of a fixed degree in the inputs.
HammingDistanceKernel Kernel over categorical inputs, based on the Hamming distance.
TanimotoKernel Kernel over molecular fingerprints, based on the Tanimoto similarity.
WassersteinKernel Kernel based on the Wasserstein distance.

Kernel

data_models.kernels.api.Kernel()

Covariance function of a Gaussian process.

A kernel \(k(\mathbf x, \mathbf x')\) gives the prior covariance between the function values at two points of the input space. It encodes the assumptions the surrogate makes about the response: how smooth it is, over what distance observations carry information, and which inputs matter at all.

AdditiveKernel

data_models.kernels.api.AdditiveKernel()

Sum of several kernels, \(k(\mathbf x, \mathbf x') = \sum_i k_i(\mathbf x, \mathbf x')\).

MultiplicativeKernel

data_models.kernels.api.MultiplicativeKernel()

Product of several kernels, \(k(\mathbf x, \mathbf x') = \prod_i k_i(\mathbf x, \mathbf x')\).

ScaleKernel

data_models.kernels.api.ScaleKernel()

Wraps another kernel with a fitted output scale, \(k(\mathbf x, \mathbf x') = \theta\, k_{\text{base}}(\mathbf x, \mathbf x')\).

The base kernel sets the shape of the covariance and this sets its magnitude, which is the variance of the noiseless signal.

MaternKernel

data_models.kernels.api.MaternKernel()

Matern kernel, a less smooth alternative to the RBF kernel.

\[ k(\mathbf x, \mathbf x') = \frac{2^{1-\nu}}{\Gamma(\nu)} \left(\sqrt{2\nu}\,\frac{d}{\ell}\right)^{\nu} K_{\nu}\!\left(\sqrt{2\nu}\,\frac{d}{\ell}\right), \qquad d = \lVert \mathbf x - \mathbf x' \rVert \]

Samples are \(\lceil \nu \rceil - 1\) times differentiable, so \(\nu\) sets how rough the response may be. Reach for this when an RBF fit looks implausibly smooth between observations.

RBFKernel

data_models.kernels.api.RBFKernel()

Radial basis function kernel, the usual default for continuous inputs.

\[ k(\mathbf x, \mathbf x') = \exp\left(-\frac{\lVert \mathbf x - \mathbf x' \rVert^2} {2\ell^2}\right) \]

Samples from this kernel are infinitely differentiable, so it assumes a very smooth response. Use MaternKernel where the response is expected to be rougher.

LinearKernel

data_models.kernels.api.LinearKernel()

Linear kernel, equivalent to Bayesian linear regression on the inputs.

\[ k(\mathbf x, \mathbf x') = v\,\mathbf x^{\top} \mathbf x' \]

PolynomialKernel

data_models.kernels.api.PolynomialKernel()

Polynomial kernel, of a fixed degree in the inputs.

\[ k(\mathbf x, \mathbf x') = (\mathbf x^{\top} \mathbf x' + c)^{p} \]

HammingDistanceKernel

data_models.kernels.api.HammingDistanceKernel()

Kernel over categorical inputs, based on the Hamming distance.

\[ k(\mathbf x, \mathbf x') = \exp\left(-\frac{d(\mathbf x, \mathbf x')}{\ell}\right) \]

where \(d\) is zero where two inputs hold the same category and one where they differ, averaged over the categorical dimensions. With ard there is one lengthscale per categorical feature, so a change of category can matter more in some features than in others. The kernel is not differentiable with respect to its inputs.

TanimotoKernel

data_models.kernels.api.TanimotoKernel()

Kernel over molecular fingerprints, based on the Tanimoto similarity.

\[ k(\mathbf x, \mathbf x') = \frac{\mathbf x^{\top}\mathbf x'} {\lVert \mathbf x \rVert^2 + \lVert \mathbf x' \rVert^2 - \mathbf x^{\top}\mathbf x'} \]

Normalizing the shared bits by the bits present in either input is what makes this the standard similarity for the sparse binary vectors a fingerprint produces.

WassersteinKernel

data_models.kernels.api.WassersteinKernel()

Kernel based on the Wasserstein distance.

It only works for 1D data that is monotonically increasing, as it is just calculating the integral of the absolute difference between two shapes. Only when both shapes are monotonically increasing, this integral is also a Wasserstein distance (https://arxiv.org/abs/2002.01878).

The shape are assumed to be discretized as a set of points. Make sure that the discretization is fine enough to capture the shape of the data.