BABEL was a joint European project under the COPERNICUS scheme comprising partners from a number of Eastern and Western European research centers. BABEL has produced a multi-language database comprising five of the most widely differing Eastern European languages: Bulgarian, Estonian, Hungarian, Polish and Romanian. 


The database has been designed and collected using the standards and protocols laid out in the European Union ESPRIT SAM project and follow the format of the EUROM 1 database; a database of eleven Central and Western European languages: Danish, Dutch, English, French, German, Italian, Norwegian, Swedish, Greek, Portuguese and Spanish.


A comparable database of spoken language has been created. For each language, as in the EUROM 1 database, there is a many-talker, few-talker and very-few-talker corpus with material minimally including prepared lists of numbers (covering the phonotactic possibilities of each language), passages and sentences. Talkers are selected equally from both sexes. Little material of this sort has been gathered for languages spoken in countries eligible for Copernicus funding, though a database of Czech exists and a small database of Bulgarian speech has been recorded and labelled by members of the present consortium using SAM protocols.


Data collection was carried out on a SAM speech workstation: this is a PC equipped with standard specified audio hardware and software. Analysis of the data comprises at least end-point labelling for all data with some more detailed labelling of other sections including phonemic/phonetic transcriptions using the SAMPA system.