I think these types of adversarial attacks are even easier to foil than that because they're specific to one particular set of weights. Even really really small changes in the training data or model could invalidate the attack if I understand correctly.
I know there has been work in generating adversarial images that work against multiple models. That kind of thing is probably only going to get better, to say nothing of particular sets of weights in a single model.