Generalisation in deep convolutional neural networks
| dc.contributor.advisor | Barnard, E | |
| dc.contributor.author | Molapo, NR | |
| dc.date.accessioned | 2026-04-21T08:01:07Z | |
| dc.date.issued | 2025 | |
| dc.description | Thesis, Doctor of Philosophy in Computer and Electronic Engineering, North-West University, 2025 | |
| dc.description.abstract | Understanding the remarkable generalization abilities of Deep Learning (DL) systems remains one of the significant scientific challenges of our time. Even though it is widely accepted that the success of Deep Neural Networks stems (at least partially) from having many hidden layers. However, the benefits of such depth are not universal. For example, in fully connected networks or Multilayer Perceptrons (MLPs), it seems likely that networks with a single hidden layer can be found to generalize as well as deeper networks. In contrast, Convolutional Neural Networks (CNNs) demonstrate a relatively robust trend of improved generalization with greater network depth. We introduce a simple experimental paradigm demonstrating the contrast between CNNs and MLPs on the MNIST, FMNIST and CIFAR-10 datasets. This paradigm demonstrates a critical distinction between modern networks and their predecessors, which is the ability of contemporary architectures to leverage deep, multi-layered structures. It consistently demonstrates improved generalization as the number of hidden layers increases. These observations, however, conflict with the classical learning theory and its key concept of the bias-variance tradeoff, which implies that over-parameterized, high-dimensional and high-complexity networks should not perform well on unseen data. Therefore, we present an alternative framework for understanding the relationship between network architecture and generalization by viewing classifiers as maps between different metric spaces. This explicit expression, represented by MLPs and CNNs with Rectified Linear Unit (ReLU) transfer functions, is derived to show that maps for CNNs are biased to be increasingly smooth with greater depth. Since such smoothness is a characteristic of real images, deeper CNNs tend to outperform shallow CNNs and MLPs. Through comparative analysis, we uncover how deeper networks develop a bias towards smoother input representations. We further argue that Wolpert's Extended Bayesian Framework (EBF) offers a promising theoretical lens to characterise this bias. While fully harnessing the insights provided by the Extended Bayesian Framework (EBF) will necessitate the creation of new theoretical and computational tools, this approach can Understanding the remarkable generalization abilities of Deep Learning (DL) systems remains one of the significant scientific challenges of our time. Even though it is widelyaccepted that the success of Deep Neural Networks stems (at least partially) from having many hidden layers. However, the benefits of such depth are not universal. For example, in fully connected networks or Multilayer Perceptrons (MLPs), it seems likely that networks with a single hidden layer can be found to generalize as well as deeper networks. In contrast, Convolutional Neural Networks (CNNs) demonstrate a relatively robust trend of improved generalization with greater network depth. We introduce a simple experimental paradigm demonstrating the contrast between CNNs and MLPs on the MNIST, FMNIST and CIFAR-10 datasets. This paradigm demonstrates a critical distinction between modern networks and their predecessors, which is the ability of contemporary architectures to leverage deep, multi-layered structures. It consistently demonstrates improved generalization as the number of hidden layers increases. These observations, however, conflict with the classical learning theory and its key concept of the bias-variance tradeoff, which implies that over-parameterized, high-dimensional and high-complexity networks should not perform well on unseen data. Therefore, we present an alternative framework for understanding the relationship between network architecture and generalization by viewing classifiers as maps between different metric spaces. This explicit expression, represented by MLPs and CNNs with Rectified Linear Unit (ReLU) transfer functions, is derived to show that maps for CNNs are biased to be increasingly smooth with greater depth. Since such smoothness is a characteristic of real images, deeper CNNs tend to outperform shallow CNNs and MLPs. Through comparative analysis, we uncover how deeper networks develop a bias towards smoother input representations. We further argue that Wolpert's Extended Bayesian Framework (EBF) offers a promising theoretical lens to characterise this bias. While fully harnessing the insights provided by the Extended Bayesian Framework (EBF) will necessitate the creation of new theoretical and computational tools, this approach can potentially deepen our understanding of generalization in Deep Learning. | |
| dc.identifier.uri | https://orcid.org/ 0000-0002-0112-9098 | |
| dc.identifier.uri | http://hdl.handle.net/10394/46654 | |
| dc.language.iso | en | |
| dc.publisher | North-West University | |
| dc.subject | Generalization | |
| dc.subject | Convolutional Neural Networks | |
| dc.subject | Deep Neural Networks | |
| dc.subject | Deep Learning | |
| dc.subject | Interpolation | |
| dc.title | Generalisation in deep convolutional neural networks | |
| dc.type | Thesis |
