Constraints on optimising encoder-only transformers for modelling sign language with human pose estimation keypoint data

Woods, Luke T.; Rana, Zeeshan A.

Constraints on optimising encoder-only transformers for modelling sign language with human pose estimation keypoint data

dc.contributor.author	Woods, Luke T.
dc.contributor.author	Rana, Zeeshan A.
dc.date.accessioned	2023-11-03T10:17:21Z
dc.date.available	2023-11-03T10:17:21Z
dc.date.issued	2023-11-02
dc.description.abstract	Supervised deep learning models can be optimised by applying regularisation techniques to reduce overfitting, which can prove difficult when fine tuning the associated hyperparameters. Not all hyperparameters are equal, and understanding the effect each hyperparameter and regularisation technique has on the performance of a given model is of paramount importance in research. We present the first comprehensive, large-scale ablation study for an encoder-only transformer to model sign language using the improved Word-level American Sign Language dataset (WLASL-alt) and human pose estimation keypoint data, with a view to put constraints on the potential to optimise the task. We measure the impact a range of model parameter regularisation and data augmentation techniques have on sign classification accuracy. We demonstrate that within the quoted uncertainties, other than ℓ2 parameter regularisation, none of the regularisation techniques we employ have an appreciable positive impact on performance, which we find to be in contradiction to results reported by other similar, albeit smaller scale, studies. We also demonstrate that the model architecture is bounded by the small dataset size for this task over finding an appropriate set of model parameter regularisation and common or basic dataset augmentation techniques. Furthermore, using the base model configuration, we report a new maximum top-1 classification accuracy of 84% on 100 signs, thereby improving on the previous benchmark result for this model architecture and dataset.	en_UK
dc.identifier.citation	Woods LT, Rana ZA. (2023) Constraints on optimising encoder-only transformers for modelling sign language with human pose estimation keypoint data. Journal of Imaging, Volume 9, Issue 11, November 2023, Article number 238	en_UK
dc.identifier.issn	2313-433X
dc.identifier.uri	https://doi.org/10.3390/jimaging9110238
dc.identifier.uri	https://dspace.lib.cranfield.ac.uk/handle/1826/20501
dc.language.iso	en	en_UK
dc.publisher	MDPI	en_UK
dc.rights	Attribution 4.0 International	*
dc.rights.uri	http://creativecommons.org/licenses/by/4.0/	*
dc.subject	sign language recognition	en_UK
dc.subject	human pose estimation	en_UK
dc.subject	classification	en_UK
dc.subject	computer vision	en_UK
dc.subject	deep learning	en_UK
dc.subject	machine learning	en_UK
dc.subject	supervised learning	en_UK
dc.subject	regularisation	en_UK
dc.subject	data augmentation	en_UK
dc.title	Constraints on optimising encoder-only transformers for modelling sign language with human pose estimation keypoint data	en_UK
dc.type	Article	en_UK

Files

Original bundle

Now showing 1 - 1 of 1

Name:: Constraints_on_optimising_encoder-only_transformers-2023.pdf
Size:: 658.25 KB
Format:: Adobe Portable Document Format
Description:

Download

License bundle

Now showing 1 - 1 of 1

Name:: license.txt
Size:: 1.63 KB
Format:: Item-specific license agreed upon to submission
Description:

Download

Collections

Staff publications (SATM)