This chapter comes in two parts. The first asks what the standard machine-learning
toolkit can do about the properties of Chapter 3: we train a network the ordinary way,
put Chapter 3’s question to it, and then look at data augmentation and adversarial
training, the two established methods for making a network more robust. The second part
asks what is missing from that toolkit once the property we care about is an arbitrary
logical specification rather than an -ball, and describes the property-driven
framework that Vehicle is built to support.
We will begin this chapter with a question: how can we train a neural network to be more robust within a desirable ?
The long tradition of robustifying neural networks in machine learning has a few methods
ready. For example, we can re-train the networks with new data that was augmented using images within the
desired -balls, or generate adversarial examples (sample images closest to the boundary of the -ball) during training. Let us briefly explore these approaches.
Humans learn by making mistakes. The same is true of neural networks. Loss functions are a way of measuring the “magnitude” of a mistake made by a neural network. For a given training input, loss functions compute a penalty proportional to the difference between the output of the network and the true output (i.e., the training label). Formally, this is written as follows:
The function represents the network, whose optimisation parameters are , and and represent the sizes of the input and output tensors respectively.
One of the simplest (yet usable) loss functions is called mean squared error, defined as:
where is the total number of data points, is the label for data point , and is the model’s predicted value for . I.e., we find the difference between the prediction and the actual value, square it, and take the average across the training dataset.
The example later in this chapter classifies images rather than predicting a
number, and for that the usual choice is cross-entropy loss. In the same
notation:
where is the number of classes, which is also the size of the output tensor,
is when data point belongs to class and
otherwise, and is the probability the network assigns to
class . Since only one term of the inner sum is non-zero, this is just the
average of of the probability given to the correct class. The penalty
therefore grows without bound as that probability approaches zero, which is what
makes cross-entropy a better fit than mean squared error when the outputs are
class probabilities rather than magnitudes.
One practical note: the networks we build below end in a plain linear layer, so
they emit unnormalised scores rather than probabilities. Both PyTorch’s
CrossEntropyLoss and TensorFlow’s SparseCategoricalCrossentropy apply the
normalisation internally, which is why no softmax appears in the model
definitions.
Models learn by iteratively tweaking their optimisation parameters with the goal of minimising the output of the loss function. The most common way to do this is using gradient descent. Formally, we wish to find the set of parameters that yields the least loss:
Let us start our first training pipeline.
Recall that Chapter 3 ended with failing to verify robustness. We used the model that was already pre-trained. Let us now try to train a similar model from scratch, in the plainest way possible, and see how robust it turns out to be.
Here, we use the Fashion MNIST dataset to train a neural network. All files used in this example can be found in the chapter-4/chapter-code directory of the tutorial repository.
First, we need some data to train on. Nothing here is specific to Vehicle, and the
same loader serves both the plain network we train in a moment and the
property-driven one later in the chapter.
SUBSET_SIZE restricts training to the first 1024 images. Plain training would happily
run over all 60,000, but evaluating a robustness property is far more expensive, so a
subset keeps the second half of this chapter runnable while still showing the effect.
Keep it an exact multiple of BATCH_SIZE; the reason becomes clear later, when we pass
n=BATCH_SIZE to the constraint loss.
There is deliberately no normalisation step.ToTensor (and dividing by 255 in
TensorFlow) leaves pixel values in , and that is where we leave them. This is
worth dwelling on, because the usual recipe for image classifiers subtracts the dataset
mean and divides by its standard deviation, and doing so here would quietly break the
chapter.
The reason is that our specification says what a valid input is:
validImage x = forall i j . 0 <= x ! i ! j <= 1, and the .idx datasets we hand to the
verifier hold pixels in exactly that range. Normalising during training would move the
network’s input space to roughly while leaving the verifier’s input space
at — so the network would be verified on inputs unlike any it was trained on.
Worse, would silently denote two different regions: a perturbation of
in normalised units is only of a raw pixel, so training would enforce a
neighbourhood almost three times tighter than the one being checked.
Keeping the data in the range the specification talks about costs us nothing here — the
network below trains perfectly well on raw pixels — and it keeps a single meaning for
across training and verification. Part II returns to this: it is a small
instance of the general difficulty of interfacing a logical specification with an
optimisation objective. One of the exercises in this chapter will give you a chance to reconstruct the whole training-verification pipeline with normalisation baked into the training.
Finally, note that the two frameworks store images differently: PyTorch’s ToTensor
produces channel-first batches of shape (N, 1, 28, 28), whereas here we append the
channel dimension last for TensorFlow, giving (N, 28, 28, 1). It makes no difference
to the network we are about to train, whose first layer flattens its input either way,
but it will matter once we start handing images to a specification.
With the data in place, the rest is unremarkable supervised training. The architecture
is small — two hidden layers — which is deliberate: a simpler network has smoother
decision boundaries, and is therefore easier both to train and to verify.
import torch
import torch.nn as nn
model = nn.Sequential(
nn.Flatten(),
nn.Linear(784,64),
nn.ReLU(),
nn.Linear(64,32),
nn.ReLU(),
nn.Linear(32,10))
optimizer = torch.optim.Adam(model.parameters(), lr=1e-3)
cross_entropy = nn.CrossEntropyLoss()for epoch inrange(num_epochs):
running_loss, correct, seen =0.0,0,0for images, labels in train_loader:
optimizer.zero_grad()
logits = model(images)
loss = cross_entropy(logits, labels)
loss.backward()
optimizer.step()
running_loss += loss.item()* labels.numel()
correct +=(logits.argmax(1)== labels).sum().item()
seen += labels.numel()print(f"Epoch: {epoch +1}, mean loss: {running_loss / seen:.4f}, "f"train accuracy: {100* correct / seen:.1f}%")
Nothing in that objective mentions robustness. The only thing being minimised is
cross-entropy, exactly as in any introductory classification example. This is the
complete script vanilla_classifier.py in the chapter code.
On 1024 images this network reaches 98% training accuracy by about epoch 75, and 99.9%
by epoch 150. Training it well past the point where the task has essentially been
learned gives:
epochs
mean loss
train accuracy
correctly classified
robust,
robust,
75
0.0785
98.1%
38/50
37/50
25/50
100
0.0413
99.5%
38/50
37/50
22/50
150
0.0103
99.9%
37/50
36/50
21/50
The first three columns are measured during training, on the 1024 images the network
learns from. The last three are measured afterwards by Vehicle, on the fifty images of
Chapter 3’s set — which are test images, held out of training. So correctly classified reports generalisation rather than memorisation, and it is a different
quantity from the training accuracy beside it.
Reading down the training columns, the mean loss almost halves between epochs 75 and 100
and falls by a factor of seven and a half by epoch 150. Reading down correctly classified over the same span: 38, then 38, then 37. All that extra optimisation bought
no improvement in generalisation whatsoever.
The two robustness columns have to be read against correctly classified, not against
50. An image the network already gets wrong cannot be robust, because
advises perturbedImage label fails at zero perturbation; such an image is counted as a
failure for a reason that has nothing to do with robustness. So correctly classified is
a ceiling on how many images could possibly be proved robust.
At the network sits exactly one image below that ceiling at every
checkpoint — 37 of 38, 37 of 38, 36 of 37. The apparent dip at epoch 150 is the ceiling
moving rather than robustness changing: one image left the correctly classified column
and took its robustness result with it.
The column tells a more interesting story. Here the checkpoints do not
agree, and they disagree in an orderly way: 25 provable images at 75 epochs, 22 at 100,
21 at 150. Training for longer has made the network less robust, and steadily so.
Note that this is not the ceiling moving. Over those same checkpoints the number of
correctly classified images goes 38, 38, 37 — down by one — while the number proved
robust falls by four.
That is worth dwelling on, because it is what one would expect. Once the training data is
fitted, further optimisation of cross-entropy cannot change which side of the decision
boundary those points fall on — so instead it draws the boundary closer to them, buying
ever greater confidence on examples that were already correct. The margin around each
point narrows. A small enough neighbourhood never notices this, which is why the
column is flat; a larger one does, and points that sat comfortably
inside the correct region begin to straddle its edge. Robustness, on this evidence, is not
merely a quantity the task objective neglects. It is one the task objective quietly
erodes, and the erosion is invisible until the question is asked at a radius wide enough
to see it. Which is precisely why robustness has to be trained for, rather than hoped for
as a by-product of fitting the data well Madry, A., Makelov, A., Schmidt, L., Tsipras, D., & Vladu, A. (2017). Towards deep learning models resistant to adversarial attacks. arXiv Preprint arXiv:1706.06083..
The figures therefore say more about the question than about the
network. At that radius, classifying an image correctly almost guarantees classifying its
whole neighbourhood correctly, so the verification is measuring generalisation more than
robustness.
What survives regardless is the shape of the objective. Once cross-entropy is satisfied,
there is nothing left in it to push the network towards anything it does not measure —
the loss falls sevenfold above while held-out accuracy does not improve at all.
Robustness is exactly such an unmeasured quantity. If we want it, it has to appear in the
objective itself.
How much room is there to improve? Before turning to methods, it is worth asking how
much of a robustness problem there is to solve, and that depends on . The
figures above mix two effects together, because thirteen of the fifty images are
misclassified and so fail at zero perturbation. Setting those aside and keeping only the
37 images the epoch-150 network classifies correctly, the same specification at two radii
gives:
provably robust
not robust
0.005
36/37
1
0.02
21/37
16
Same network, same images; only the size of the neighbourhood differs. A four-fold
increase in the radius turns one failure into sixteen, so the near perfect robustness at
was less a property of the network than of the question we asked it:
the neighbourhoods were small enough that almost nothing could go wrong inside them.
This sets the terms for the rest of the chapter. To ask whether adding robustness to the
training objective helps, we need a radius at which the network is genuinely vulnerable.
At there is one image to win back; at there are
sixteen.
We will first consider traditional machine learning methods that were introduced to address the problem of adversarial robustness.
Machine learning has a long tradition of making networks more robust, and two methods
dominate it. Both work by training on extra inputs drawn from around each data point;
they differ in how those inputs are chosen.
Data Augmentation works by generating additional data within the -balls of the original training data points, usually done using methods such as rotation, cropping, flipping, random sampling, etc. The augmented data points are assigned the same label as the original ones they were augmented from. We can then use our usual training methods with this augmented dataset with the hope that it will improve the network’s average-case robustness Shorten, C., & Khoshgoftaar, T. M. (2019). A survey on image data augmentation for deep learning. J. Big Data, 6(60)..
Unfortunately, this approach has its problems. Firstly, if our original sampled data point is already very close to the decision boundary, there is a chance that an augmented data point will actually lie on the wrong side, even though it is still within the -ball. This means it will have been assigned the wrong label:
In the case where two data points’ -balls overlap, there is a chance we generate two new data points with the same position in the input space. Furthermore, if the two original data points lie both close to (and on opposite sides of) the decision boundary, the augmented data points may have different labels, despite occupying the same location in the input space:
These inconsistencies mean this approach is generally unviable for network robustification.
Adversarial TrainingMadry, A., Makelov, A., Schmidt, L., Tsipras, D., & Vladu, A. (2017). Towards deep learning models resistant to adversarial attacks. arXiv Preprint arXiv:1706.06083. also involves generating new data to train the network, but unlike data augmentation where perturbations are sampled randomly, adversarial training aims to find the worst-case perturbation within -distance to a data point from the training dataset. Whilst data augmentation can be done using worst-case examples, it is still subtly different to adversarial training. Most notably, adversarial training is a process integrated into the network’s training loop, so perturbations are regenerated at every iteration. This means the worst-case examples will always be worst case, which is not true for data augmentation, as after a certain number of iterations the network will have learnt to account for these examples. Pictorially, this amounts to looking at gradients, while not drawing any firm lines:
Formally, adversarial training uses a variant of gradient descent, called projected gradient descent, to maximise loss in order to find worst-case perturbations. We ensure that the perturbation still lies within the -ball of the original data point by projecting those perturbations that escape the -ball back inside. Our new training objective, due to Madry et al. Madry, A., Makelov, A., Schmidt, L., Tsipras, D., & Vladu, A. (2017). Towards deep learning models resistant to adversarial attacks. arXiv Preprint arXiv:1706.06083., becomes:
In other words, we want to find the perturbation that is within -distance of the data point that produces the largest loss (the worst-case perturbation). Then, we aim to find the optimisation parameters which minimises this loss value.
Adversarial training works, and for -ball robustness on image classifiers it is
the standard answer. But it is worth being precise about which property it improves,
because the answer is narrower than it first appears — and this is where Part II begins.
Recall how Chapter 3 stated robustness:
.
That leaves undefined, and different ways of filling it in give genuinely
different properties. Casadio et al. Casadio, M., Komendantskaya, E., Daggitt, M. L., Kokke, W., Katz, G., Amir, G., & Refaeli, I. (2022). Neural Network Robustness as a Verification Property: A Principled Case Study. Computer Aided Verification - 34th International Conference, CAV 2022, Haifa, Israel, August 7-10, 2022, Proceedings, Part I, 13371, 219–231. set them side by side:
Training method
Definition of it optimises
Property
Data augmentation
classification robustness
DL2 training
strong classification robustness
Adversarial training
standard robustness
Lipschitz continuity
Lipschitz robustness
DL2 training refers to the method of Fischer et al. Fischer, M., Balunovic, M., Drachsler-Cohen, D., Gehr, T., Zhang, C., & Vechev, M. T. (2019). DL2: Training and Querying Neural Networks with Logic. Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA, 97, 1931–1941. http://proceedings.mlr.press/v97/fischer19a.html; is the
confidence margin it requires of the advised class.
Projected gradient descent optimises exactly one row of this table: it minimises how far
the output can move within the ball, which is standard robustness. Chapter 3’s
specification, meanwhile, asks that the advised label stays the same, which is
classification robustness — a different row.
The two are not in an implication relation in either direction. So it is entirely
possible, and in the literature common, to train for one property, verify another, and
report the result as though a single notion of robustness had been improved. Getting more
of what you optimised while gaining nothing you verified is not a subtle failure mode; it
is the expected outcome of a mismatch nobody wrote down.
This is the first of three difficulties that motivate the rest of the chapter.
Part I ended on a mismatch: adversarial training optimises standard robustness while our
specification asks for classification robustness. That is one instance of a general
problem. The machine-learning toolkit was built for one property, on one kind of input
region, in one application domain; a specification language lets us write down far more
than that. This part follows the framework of Flinkow et al. Flinkow, T., Casadio, M., Kessler, C., Monahan, R., & Komendantskaya, E. (2025). A General Framework for Property-Driven Machine Learning., which
sets out what has to change, and why Vehicle is organised the way it is.
Three difficulties stand between the standard recipe and training for arbitrary
specifications. We take them in turn.
The table at the end of Part I is the first difficulty in miniature. Interpreting a
logical specification as an optimisation objective is done by hand, informally, and it is
easy to get wrong — and when it goes wrong nothing complains. Training proceeds, the
loss falls, and the verifier reports no improvement, because the quantity being minimised
was never the quantity being checked.
The consequence Casadio et al. Casadio, M., Komendantskaya, E., Daggitt, M. L., Kokke, W., Katz, G., Amir, G., & Refaeli, I. (2022). Neural Network Robustness as a Verification Property: A Principled Case Study. Computer Aided Verification - 34th International Conference, CAV 2022, Haifa, Israel, August 7-10, 2022, Proceedings, Part I, 13371, 219–231. draw is worth stating plainly: one kind
of robustness does not imply another, so optimising for one can achieve very little in
verification success rates for another. What is needed is not a better hand-translation
but a systematic one — a single source of truth from which both the verification query
and the training objective are derived. That is exactly what a specification language can
provide, and it is the reason training belongs inside the verification toolchain rather
than beside it.
The second difficulty is the shape of the input region. Adversarial training assumes an
-norm ball around a data point,
which suits images, where a small perturbation of every pixel is a meaningful notion of
“nearby”. It suits other domains badly.
In natural language processing the input space is discrete, and an -ball around
a sentence contains no sentences — the region that matters is the set of semantically
similar sentences, which is not a ball around anything. As Chapter 2, has shown, in cyber-physical systems the
input space is low-dimensional and the interesting regions are named by the
specification itself: “intruder near and approaching from the left” is a constraint on
five variables with different units and ranges, not a ball.
The generalisation is to a hyper-rectangle, an independent interval per dimension:
We can always “draw” such a hyper-rectangle, assuming that we have only linear constraints on the input variable. And so far, all examples we cared about were given in that format.
Recall, for example, this constraint from Chapter 2:
These constriants give rise to the following hyper-rectangle in five-dimensional space:
Note that here the upper bounds for , and are those listed as maximum input values in the ACAS Xu specification in Chapter 2. Every -ball is a hyper-rectangle with and
, so nothing is lost, and regions that no ball can express become
available. The training objective generalises by substitution — where adversarial
training maximises over , we maximise
over .
The third difficulty is on the output side. Classification robustness assumes the goal is to
keep the predicted class fixed. Many specifications say something less prescriptive.
ACAS Xu’s third property, which Chapter 2 verified, asks that if the intruder is directly
ahead and closing, then the score for clear-of-conflict will not be minimal. It does not
say which advisory the network should give — only that one of the four alternatives must
outrank one particular option. Written out, the conclusion is a disjunction:
There is no label to hold fixed here, so there is nothing for the standard recipe to
maximise. What is needed is a way to take an arbitrary specification and produce a
differentiable loss that measures how far the network is from satisfying
it. Such translations are called differentiable logics.
A differentiable logic replaces the Boolean connectives with real-valued, differentiable
ones, so that a formula evaluates to a number that can be minimised rather than a truth
value that cannot. A very small example over a toy language conveys the idea:
An atom is translated into the margin by which it holds or fails, and the connectives
combine margins. Several such logics exist, and they differ in ways that matter for
optimisation Fischer, M., Balunovic, M., Drachsler-Cohen, D., Gehr, T., Zhang, C., & Vechev, M. T. (2019). DL2: Training and Querying Neural Networks with Logic. Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA, 97, 1931–1941. http://proceedings.mlr.press/v97/fischer19a.htmlSlusarz, N., Komendantskaya, E., Daggitt, M. L., Stewart, R. J., & Stark, K. (2023). Logic of Differentiable Logics: Towards a Uniform Semantics of DL. LPAR 2023: Proceedings of 24th International Conference on Logic for Programming, Artificial Intelligence and Reasoning, Manizales, Colombia, 4-9th June 2023, 94, 473–493.van Krieken, E., Acar, E., & van Harmelen, F. (2022). Analyzing Differentiable Fuzzy Logic Operators. Artif. Intell., 302, 103602.. The logic we adopt is quantitative linear logic (QLL), due to Capucci et al.
Capucci, M., Atkey, R., Grellois, C., & Komendantskaya, E. (2026). Quantitative Linear Logic.:
Negation:
Conjunction:
Disjunction:
Implication:
where is the hardness degree. As the connectives
converge on and , their traditional counterparts; smaller gives a
smoother surface whose gradients reach further. Conjunction and disjunction are the
familiar log-sum-exp softening of and , which is what makes them
differentiable everywhere.
Note that the truth direction is reversed: is the top and is the
bottom. This is usual in the differentiable logic literature, DL2 Fischer, M., Balunovic, M., Drachsler-Cohen, D., Gehr, T., Zhang, C., & Vechev, M. T. (2019). DL2: Training and Querying Neural Networks with Logic. Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA, 97, 1931–1941. http://proceedings.mlr.press/v97/fischer19a.html
included, because the value measures how far a formula is from being satisfied — the
less error there is, the more true the formula is. It also means that a loss built this
way is unbounded below, which matters when it is combined with a task loss.
The three pieces now assemble. Standard training minimises the expected task loss over
the data:
Adversarial training minimises it against the worst case in an -ball, and
Problem 2 replaces that ball with a hyper-rectangle:
Problem 3 adds the specification itself as a second term, translated by the
differentiable logic, and balances the two:
This single objective has the earlier methods as special cases. Setting
recovers adversarial training over a general region. Taking to be the
-cube around and to be
recovers standard
robustness — the row of Part I’s table that adversarial training optimises. And a
specification with no input constraint at all, , is
handled by letting be the domain of the input, typically the normalisation
bounds , .
Both extremes of are worth understanding. At the specification is
ignored. At the task is ignored, and since a constant network satisfies most
robustness properties perfectly, the optimiser is free to discard the classifier
altogether — the constraint term alone does not distinguish a useful constant from a
useless one. The blend is not a convenience; it is what rules out the degenerate solution.
Vehicle’s answer to Problem 1 is to derive both artefacts from one source. The same
.vcl specification that Chapter 3 compiled into verification queries can be compiled
into a loss function, so the property being trained for and the property being verified
are the same text. Nothing is hand-translated, and the two cannot drift apart.
The interface mirrors the objective above. load_specification takes the specification
and a differentiable logic and returns the named properties, each as a callable that
evaluates for a batch; alpha in the training loop is the that
balances task loss against constraint loss.
The code below selects VehicleDifferentiableLogic, Vehicle’s built-in default. The
Capucci logic of the previous section can be declared directly in the specification as a
DifferentiableTensorLogic and selected by name with
vcl.CustomDifferentiableLogic("qllAdditive"); a worked example lives alongside the
chapter code.
A note on the current release. The sections above describe the framework and the
Vehicle interface, and both are stable. We do not, however, present trained-and-verified
results here. In Vehicle 0.27.1 the loss compiled from a forall quantifier does not
behave as the objective above requires: widening the input region, or searching it
harder, makes the compiled loss report the property as better satisfied rather than
worse. Since the framework depends on that inner maximisation being a genuine worst case,
we have deferred the experimental half of this chapter until the behaviour is resolved
upstream, rather than report numbers we cannot stand behind. Readers can reproduce the
diagnostic themselves by evaluating a specification’s loss at several values of
and watching which way it moves.
Next, we will load our Vehicle specification and define our constraint loss function:
import vehicle_lang as vcl
from vehicle_lang.loss import pytorch as loss_pt
spec = loss_pt.load_specification("fmnist-robustness.vcl",
logic=vcl.VehicleDifferentiableLogic(),)
constraint_loss_fn = spec["robust"]
import vehicle_lang as vcl
from vehicle_lang.loss import tensorflow as loss_tf
spec = loss_tf.load_specification("fmnist-robustness.vcl",
logic=vcl.VehicleDifferentiableLogic(),)
constraint_loss_fn = spec["robust"]
The first parameter to the load_specification function is the path to the Vehicle specification. The second parameter defines which logic to use – this is optional, and defaults to DL2. We define which property from the specification to use as our constraint loss function by accessing it by name on the specification object.
Next, we will define a simple model and training procedure:
Note that the network callable must match the type of the network declared in the Vehicle specification. The alpha parameter can be used to tweak the weighting of task loss vs. constraint loss, which are blended together in total_loss.
The n=BATCH_SIZE argument is the reason SUBSET_SIZE had to divide evenly by BATCH_SIZE. The specification declares its data set size as an inferred parameter, n, and here each batch plays the role of the data set, so n must be exactly the number of images in the batch. Were the last batch of an epoch short, the value of n would no longer describe the images being passed alongside it. We can now export this trained model to verify it using Vehicle, with the hope that it is more robust as a result of training with constraint loss. Model exportation can be done like so:
import torch.onnx
model.eval()
input_tensor = torch.randn(1,1,28,28)
torch.onnx.export(
model,
input_tensor,"classifier.onnx",# file name
external_data=False,# required for Marabou verification)
Exporting an ONNX file in PyTorch works by tracing, which runs the model with an arbitrary input and records each operation. Hence, we provide the model with a randomly generated input tensor. At the time of writing, Marabou does not support external data locations, so we require that external_data=False. Recent versions of PyTorch route torch.onnx.export through a new exporter that additionally requires the onnxscript package, so install that alongside torch if the export reports it missing.
model.export("classifier")
This saves the model at the specified directory. To convert this to ONNX format, we can run the following command (this will require installing tf2onnx):
Download the required materials (or produce them yourself) and repeat the steps described in this chapter. All code used in this chapter is available from the tutorial repository.
Use a Vehicle specification (either the one provided, or your own) to verify a property (e.g., robustness) of a network trained with a logical loss function, and compare this to one trained without a logical loss function. Which is more robust, which has the better task accuracy, and why?
Try various combinations of task loss functions, constraint loss functions, and alpha values. How do these affect each other? Is there a combination that makes the network more robust? Is there a combination that makes the network more accurate? What happens when you use multiple constraint loss functions simultaneously?
Finally, try creating your own model from scratch and repeat the experiments and comparisons described above. Explore the relationship between how complex a model is and to what degree it can satisfy robustness, and the effect robustness training can have on this.
Hint: a simple model is worse at spotting the difference between two different images. Does this make it more or less likely to be robust?