Artificial intelligence
A dataset is an ordered collection of data used to train, tune and test a machine learning model.
It is usually split into three parts: one for training, one for tuning the model and one for measuring its result on data it has never seen. The quality of the dataset decides the quality of the model: incomplete or skewed data produce wrong predictions and algorithmic bias. The case that showed how much this matters is ImageNet, the large archive of labelled images built from 2009 by the group of the researcher Fei-Fei Li at Stanford: on that data, in 2012, a neural network demonstrated the power of deep learning. In Europe the AI Act requires the datasets of high-risk systems to be relevant, representative and as free of errors as possible.
An example
Three years of customer purchase history, flagged with who stopped buying, used to teach a model to recognise who is about to leave.
What are your business challenges?
Tell us about your priorities and the objectives you want to reach.
You will receive a free, targeted answer within one working day.