Filters
total: 474
Search results for: DATASET QUALITY
-
Towards semantic-rich word embeddings
PublicationIn recent years, word embeddings have been shown to improve the performance in NLP tasks such as syntactic parsing or sentiment analysis. While useful, they are problematic in representing ambiguous words with multiple meanings, since they keep a single representation for each word in the vocabulary. Constructing separate embeddings for meanings of ambiguous words could be useful for solving the Word Sense Disambiguation (WSD)...
-
Simultaneous grouping and ranking with combination of SOM and TOPSIS for selection of preferable analytical procedure for furan determination in food
PublicationNovel methodology for grouping and ranking with application of self-organizing maps and multicriteria decision analysis is presented. The dataset consists of 22 objects that are analytical procedures applied to furan determination in food samples. They are described by 10 variables, referred to their analytical performance, environmental and economic aspects. Multivariate statistics analysis allows to limit the amount of input...
-
Global Value Chains and Wages: Multi-Country Evidence from Linked Worker-Industry Data
PublicationThis paper uses a multi-country microeconomic setting to contribute to the literature on the nexus between production fragmentation and wages. Exploiting a rich dataset on over 110,000 workers from nine Eastern and Western European countries and the United States, we study the relationship between individual workers’ wages and industry ties into global value chains (GVCs). We find an inverse (but weak) relationship between the...
-
Bi-GRU-APSO: Bi-Directional Gated Recurrent Unit with Adaptive Particle Swarm Optimization Algorithm for Sales Forecasting in Multi-Channel Retail
PublicationIn the present scenario, retail sales forecasting has a great significance in E-commerce companies. The precise retail sales forecasting enhances the business decision making, storage management, and product sales. Inaccurate retail sales forecasting can decrease customer satisfaction, inventory shortages, product backlog, and unsatisfied customer demands. In order to obtain a better retail sales forecasting, deep learning models...
-
Ontological Modeling for Contextual Data Describing Signals Obtained from Electrodermal Activity for Emotion Recognition and Analysis
PublicationMost of the research in the field of emotion recognition is based on datasets that contain data obtained during affective computing experiments. However, each dataset is described by different metadata, stored in various structures and formats. This research can be counted among those whose aim is to provide a structural and semantic pattern for affective computing datasets, which is an important step to solve the problem of data...
-
Split-beam echosounder data from Puck Bay autumn 2018 Part II
Open Research DataThe acoustic data was collected in 2018, in the Bay of Puck, in the seasons: autumn. Data was collected during the day and night. Three split-beam echosounders with frequencies of 38 kHz, 120 kHz and 333 kHz were used to collect the data. The data was collected at a designated study area not far from the city of Hel and on the route Hel - Gdynia while...
-
Automatic music genre classification based on musical instrument track separation / Automatyczna klasyfikacja gatunku muzycznego wykorzystująca algorytm separacji dźwięku instrumentó muzycznych
PublicationThe aim of this article is to investigate whether separating music tracks at the pre-processing phase and extending feature vector by parameters related to the specific musical instruments that are characteristic for the given musical genre allow for efficient automatic musical genre classification in case of database containing thousands of music excerpts and a dozen of genres. Results of extensive experiments show that the approach...
-
Data from the Survey on Gdańsk University of Technology Graduates’ Professional Careers
PublicationThe dataset titled Data from the survey on Gdańsk University of Technology graduates’ professional careers includes data from a survey of Gdańsk University of Technology (Gdańsk Tech) graduates’ professional careers. The survey was conducted in 2017, two years after the respondents obtained graduate status. The research sample included 2553 respondents. The study concerned, i.a. the percentage of people working among graduates...
-
Bus bays inventory using a terrestrial laser scanning system
PublicationThis article presents the use of laser scanning technology for the assessment of bus bay geo-location. Ground laser scanning is an effective tool for collecting three-dimensional data. Moreover, the analysis of a point cloud dataset can be a source of a lot of information. The authors have outlined an innovative use of data collection and analysis using the TLS regarding information on the flatness of bus bays. The results were...
-
TOWARDS EXPLAINABLE CLASSIFIERS USING THE COUNTERFACTUAL APPROACH - GLOBAL EXPLANATIONS FOR DISCOVERING BIAS IN DATA
PublicationThe paper proposes summarized attribution-based post-hoc explanations for the detection and identification of bias in data. A global explanation is proposed, and a step-by-step framework on how to detect and test bias is introduced. Since removing unwanted bias is often a complicated and tremendous task, it is automatically inserted, instead. Then, the bias is evaluated with the proposed counterfactual approach. The obtained results...
-
Neural Network Subgraphs Correlation with Trained Model Accuracy
PublicationNeural Architecture Search (NAS) is a computationally demanding process of finding optimal neural network architecture for a given task. Conceptually, NAS comprises applying a search strategy on a predefined search space accompanied by a performance evaluation method. The design of search space alone is expected to substantially impact NAS efficiency. We consider neural networks as graphs and find a correlation between the presence...
-
Global Value Chains and Wages: International Evidence from Linked Worker-Industry Data
PublicationUsing a rich dataset on over 110,000 workers from nine European countries and the USA we study the wage response to industry dependence on foreign value added. We estimate a Mincerian wage model augmented with an input-output interindustry linkages measure accounting for task heterogeneity across workers. Low and mediumeducated workers and those performing routine tasks experience (little) wage decline due to major dependency of...
-
Data from the Survey on Entrepreneurs’ Opinions on Factors Determining the Employment of the Gdańsk University of Technology Graduates
PublicationThe dataset includes data from a survey on factors determining the employment of the Gdańsk University of Technology (Gdańsk Tech) graduates’ in the opinion of entrepreneurs. The survey was conducted in 2017. The research sample included 102 respondents representing various firms from the Pomeranian Voivodeship, Poland. The study concerned i.a. factors determining the decision to hire a candidate, methods of recruiting employees,...
-
Cascade Object Detection and Remote Sensing Object Detection Method Based on Trainable Activation Function
PublicationObject detection is an important process in surveillance system to locate objects and it is considered as major application in computer vision. The Convolution Neural Network (CNN) based models have been developed by many researchers for object detection to achieve higher performance. However, existing models have some limitations such as overfitting problem and lower efficiency in small object detection. Object detection in remote...
-
Mobilenet-V2 Enhanced Parkinson's Disease Prediction with Hybrid Data Integration
PublicationThis study investigates the role of deep learning models, particularly MobileNet-v2, in Parkinson's Disease (PD) detection through handwriting spiral analysis. Handwriting difficulties often signal early signs of PD, necessitating early detection tools due to potential impacts on patients' work capacities. The study utilizes a three-fold approach, including data augmentation, algorithm development for simulated PD image datasets,...
-
Acquisition and indexing of RGB-D recordings for facial expressions and emotion recognition
PublicationIn this paper KinectRecorder comprehensive tool is described which provides for convenient and fast acquisition, indexing and storing of RGB-D video streams from Microsoft Kinect sensor. The application is especially useful as a supporting tool for creation of fully indexed databases of facial expressions and emotions that can be further used for learning and testing of emotion recognition algorithms for affect-aware applications....
-
Tweet you right back: Follower anxiety predicts leader anxiety in social media interactions during the SARS-CoV-2 pandemic
PublicationRecent research has shown that organizational leaders’ tweets can influence employee anxiety. In this study, we turn the table and examine whether the same can be said about followers’ tweets. Based on emotional contagion and a dataset of 108 leaders and 178 followers across 50 organizations, we infer and track state- and trait-anxiety scores of participants over 316 days, including pre- and post the onset of the SARS-CoV-2 pandemic...
-
Tribological Properties of Thermoplastic Materials Formed by 3D Printing by FDM Process
PublicationThe dataset entitled 3D printed ABS thermoplastic vs. steel. Dry sliding wear test in constant load & velocity ring on flat configuration. Test parameters: print layer thickness and orientation. Test symbol: 019_h_4 contains: the time base (expressed in seconds and minutes), the friction torque for sliding friction, rotational velocity of the counter – specimen (velocity of sliding), friction coefficient, load in the friction contact...
-
Integrating Statistical and Machine‐Learning Approach for Meta‐Analysis of Bisphenol A‐Exposure Datasets Reveals Effects on Mouse Gene Expression within Pathways of Apoptosis and Cell Survival
PublicationBisphenols are important environmental pollutants that are extensively studied due to different detrimental effects, while the molecular mechanisms behind these effects are less well understood. Like other environmental pollutants, bisphenols are being tested in various experimental models, creating large expression datasets found in open access storage. The meta‐analysis of such datasets is, however, very complicated for various...
-
Exploring music listening patterns: an online survey
PublicationAn online survey was carried out to explore how respondents listen to music recordings. It was anticipated that the listener’s preferences would be influenced by various factors, such as age, music genre, the contexts in which they listen, and their favored methods of music consumption. Consequently, the data were collected to analyze these relationships. The survey, structured as a web application, encompassed 23 questions,...
-
Data-driven Models for Predicting Compressive Strength of 3D-printed Fiber-Reinforced Concrete using Interpretable Machine Learning Algorithms
Publication3D printing technology is growing swiftly in the construction sector due to its numerous benefits, such as intricate designs, quicker construction, waste reduction, environmental friendliness, cost savings, and enhanced safety. Nevertheless, optimizing the concrete mix for 3D printing is a challenging task due to the numerous factors involved, requiring extensive experimentation. Therefore, this study used three machine learning...
-
CPLFD-GDPT5: High-resolution gridded daily precipitation and temperature data set for two largest Polish river basins
PublicationThe CHASE-PL (Climate change impact assessment for selected sectors in Poland) Forcing Data–Gridded Daily Precipitation & Temperature Dataset–5 km (CPLFD-GDPT5) consists of 1951–2013 daily minimum and maximum air temperatures and precipitation totals interpolated onto a 5 km grid based on daily meteorological observations from the Institute of Meteorology and Water Management (IMGW-PIB; Polish stations), Deutscher Wetterdienst...
-
Real-Time Facial Features Detection from Low Resolution Thermal Images with Deep Classification Models
PublicationDeep networks have already shown a spectacular success for object classification and detection for various applications from everyday use cases to advanced medical problems. The main advantage of the classification models over the detection models is less time and effort needed for dataset preparation, because classification networks do not require bounding box annotations, but labels at the image level only. Yet, after passing...
-
AVHRR Level1CD covering Baltic Sea area year 2006
Open Research DataThe product level is the NOAA AVHRR Level 1C that is result of processing the AVHRR data from the HRPT stream based on ancillary information like sensing geometry and calibration data. Then converted into geophysical variables: top-of-the atmosphere (TOA) albedo or brightness temperature. Additionally, information like geolocation has been added. Other...
-
AVHRR Level1CD covering Baltic Sea area year 2010
Open Research DataThe product level is the NOAA AVHRR Level 1C that is result of processing the AVHRR data from the HRPT stream based on ancillary information like sensing geometry and calibration data. Then converted into geophysical variables: top-of-the atmosphere (TOA) albedo or brightness temperature. Additionally, information like geolocation has been added. Other...
-
AVHRR Level1CD covering Baltic Sea area year 2007
Open Research DataThe product level is the NOAA AVHRR Level 1C that is result of processing the AVHRR data from the HRPT stream based on ancillary information like sensing geometry and calibration data. Then converted into geophysical variables: top-of-the atmosphere (TOA) albedo or brightness temperature. Additionally, information like geolocation has been added. Other...
-
AVHRR Level1CD covering Baltic Sea area year 2011
Open Research DataThe product level is the NOAA AVHRR Level 1C that is result of processing the AVHRR data from the HRPT stream based on ancillary information like sensing geometry and calibration data. Then converted into geophysical variables: top-of-the atmosphere (TOA) albedo or brightness temperature. Additionally, information like geolocation has been added. Other...
-
AVHRR Level1CD covering Baltic Sea area year 2012
Open Research DataThe product level is the NOAA AVHRR Level 1C that is result of processing the AVHRR data from the HRPT stream based on ancillary information like sensing geometry and calibration data. Then converted into geophysical variables: top-of-the atmosphere (TOA) albedo or brightness temperature. Additionally, information like geolocation has been added. Other...
-
AVHRR Level1CD covering Baltic Sea area year 2008
Open Research DataThe product level is the NOAA AVHRR Level 1C that is result of processing the AVHRR data from the HRPT stream based on ancillary information like sensing geometry and calibration data. Then converted into geophysical variables: top-of-the atmosphere (TOA) albedo or brightness temperature. Additionally, information like geolocation has been added. Other...
-
AVHRR Level1CD covering Baltic Sea area year 2009
Open Research DataThe product level is the NOAA AVHRR Level 1C that is result of processing the AVHRR data from the HRPT stream based on ancillary information like sensing geometry and calibration data. Then converted into geophysical variables: top-of-the atmosphere (TOA) albedo or brightness temperature. Additionally, information like geolocation has been added. Other...
-
Measurements of ultrasonic bulk and guided wave propagation in additively manufactured cubes and plates obtained by ultrasonic pulse velocity analyzer and scanning laser vibrometry
Open Research DataThe DataSet contains the results of measurements of ultrasonic wave propagation in additively manufactured samples made of polylactic acid (PLA). Three types of raster angles in two consecutive layers were assumed: 0°/90° (#1), 45°/-45° (#2) and 90°/90° (#3) with respect to the x-axis. For each printing variants a cubic sample (#C1-3) with dimensions...
-
A Bayesian regularization-backpropagation neural network model for peeling computations
PublicationA Bayesian regularization-backpropagation neural network (BRBPNN) model is employed to predict some aspects of the gecko spatula peeling, viz. the variation of the maximum normal and tangential pull-off forces and the resultant force angle at detachment with the peeling angle. K-fold cross validation is used to improve the effectiveness of the model. The input data is taken from finite element (FE) peeling results. The neural network...
-
Occurrence of Cyanobacteria in the Gulf of Gdańsk (2008–2009)
PublicationBlooms of cyanobacteria develop each summer in the Baltic Sea. Collecting complete data on this phenomenon is helpful in understanding the changes taking place in the Baltic Sea and forecasting the occurrence of these phenomena in the future. This dataset includes unpublished information about the occurrence of cyanobacteria in the Gulf of Gdańsk (Southern Baltic) in 2008 and 2009. The presented data combines basic physic-ochemical...
-
Driver fatigue detection method based on facial image analysis
PublicationNowadays, ensuring road safety is a crucial issue that demands continuous development and measures to minimize the risk of accidents. This paper presents the development of a driver fatigue detection method based on the analysis of facial images. To monitor the driver's condition in real-time, a video camera was used. The method of detection is based on analyzing facial features related to the mouth area and eyes, such as...
-
Neural network model of ship magnetic signature for different measurement depths
PublicationThis paper presents the development of a model of a corvette-type ship’s magnetic signature using an artificial neural network (ANN). The capabilities of ANNs to learn complex relationships between the vessel’s characteristics and the magnetic field at different depths are proposed as an alternative to a multi-dipole model. A training dataset, consisting of signatures prepared in finite element method (FEM) environment Simulia...
-
Preeclampsia Risk Prediction Using Machine Learning Methods Trained on Synthetic Data
PublicationThis paper describes a research study that investigates the use of machine learning algorithms on synthetic data to classify the risk of developing preeclampsia by pregnant women. Synthetic datasets were generated based on parameter distributions from three real patient studies. Four models were compared: XGBoost, Support Vector Machine (SVM), Random Forest, and Explainable Boosting Machines (EBM). The study found that the XGBoost...
-
Photos and rendered images of LEGO bricks
PublicationThe paper describes a collection of datasets containing both LEGO brick renders and real photos. The datasets contain around 155,000 photos and nearly 1,500,000 renders. The renders aim to simulate real-life photos of LEGO bricks allowing faster creation of extensive datasets. The datasets are publicly available via the Gdansk University of Technology “Most Wiedzy” institutional repository. The source files of all tools used during...
-
How Specific Can We Be with k-NN Classifier?
PublicationThis paper discusses the possibility of designing a two stage classifier for large-scale hierarchical and multilabel text classification task, that will be a compromise between two common approaches to this task. First of it is called big-bang, where there is only one classifier that aims to do all the job at once. Top-down approach is the second popular option, in which at each node of categories’ hierarchy, there is a flat classifier...
-
Instance segmentation of stack composed of unknown objects
PublicationThe article reviews neural network architectures designed for the segmentation task. It focuses mainly on instance segmentation of stacked objects. The main assumption is that segmentation is based on a color image with an additional depth layer. The paper also introduces the Stacked Bricks Dataset based on three cameras: RealSense L515, ZED2, and a synthetic one. Selected architectures: DeepLab, Mask RCNN, DEtection TRansformer,...
-
The Belt and Road Initiative and export variety: 1996–2019
PublicationThis study examines the association between the Belt and Road Initiative (BRI) and export variety (EV). We propose three hypotheses on how BRI may foster export markets (destinations) or export product lines. The estimates are based on a dataset constructed specifically for this analysis, covering 183 countries and linked with trade data from 1996 to 2019. We apply the instrumental variable (IV) approach in regressions for covering the...
-
LSA Is not Dead: Improving Results of Domain-Specific Information Retrieval System Using Stack Overflow Questions Tags
PublicationThe paper presents the approach to using tags from Stack Overflow questions as a data source in the process of building domain-specific unsupervised term embeddings. Using a huge dataset of Stack Overflow posts, our solution employs the LSA algorithm to learn latent representations of information technology terms. The paper also presents the Teamy.ai system, currently developed by Scalac company, which serves as a platform that...
-
Expectation-Maximization Model for Substitution of Missing Values Characterizing Greenness of Organic Solvents
PublicationOrganic solvents are ubiquitous in chemical laboratories and the Green Chemistry trend forces their detailed assessments in terms of greenness. Unfortunately, some of them are not fully characterized, especially in terms of toxicological endpoints that are time consuming and expensive to be determined. Missing values in the datasets are serious obstacles, as they prevent the full greenness characterization of chemicals. A featured...
-
Analysis of results of large-scale multimodal biometric identity verification experiment
PublicationAn analysis of a large set of biometric data obtained during the enrolment and the verification phase in an experimental biometric system installed in bank branches is presented. Subjective opinions of bank clients and of bank tellers were also surveyed concerning the studied biometric methods in order to discover and to explore relations emerging from the obtained multimodal dataset. First, data acquisition and identity verification...
-
Hasse diagram as a green analytical metrics tool: ranking of methods for benzo[a]pyrene determination in sediments
PublicationThis study presents an application of the Hasse diagram technique (HDT) as the assessment tool to select the most appropriate analytical procedures according to their greenness or the best analytical performance. The dataset consists of analytical procedures for benzo[a]pyrene determination in sediment samples, which were described by 11 variables concerning their greenness and analytical performance. Two analyses with the HDT...
-
Residual MobileNets
PublicationAs modern convolutional neural networks become increasingly deeper, they also become slower and require high computational resources beyond the capabilities of many mobile and embedded platforms. To address this challenge, much of the recent research has focused on reducing the model size and computational complexity. In this paper, we propose a novel residual depth-separable convolution block, which is an improvement of the basic...
-
Study on the student's perception of marketing activities undertaken by Gdansk and Mangalore
Open Research DataNowadays, the urban areas very often attract the poor and the unemployed, leading to the creation of neighbourhoods of poverty (slums) and other economic and social problems. All over the World cities include sustainable goals in their development strategies but the question is, whether the city development strategies foresee activities devoted also...
-
Results of calibration of the piezoelectric scanner using the probe TGQ1
Open Research DataTeaching file. Results of calibration of the piezoelectric scanner using the probe TGQ1. Scanning in contact mode. NTEGRA Prima (NT-MDT) device. CSG probe 10.
-
The surface of a fragment of the structure of an integrated circuit in the semi-contact mode.
Open Research DataThe surface of a fragment of the structure of an integrated circuit. Topographic measurements in the semi-contact mode. NTEGRA Prima (NT-MDT) device. NSG 01 probe.
-
Car wash water reuse
Open Research DataThe resistance of automotive coatings to washing water recovered in 50% and 70% from wastewater generated at car wash was tested. The wastewater was purified in the ultrafiltration process using tubular polyvinylidene fluoride membranes (100 and 200 kDa) manufactured by PCI company. The membranes retained oil contamination and over 50% of surfactants....
-
Vehicle detector training with minimal supervision
PublicationRecently many efficient object detectors based on convolutional neural networks (CNN) have been developed and they achieved impressive performance on many computer vision tasks. However, in order to achieve practical results, CNNs require really large annotated datasets for training. While many such databases are available, many of them can only be used for research purposes. Also some problems exist where such datasets are not...