TY - GEN
T1 - Smoking Cessation Recruitment Analysis
T2 - 2017 IEEE International Conference on Big Knowledge, ICBK 2017
AU - Li, Wei
AU - Cui, Xiaohui
AU - Amaral, Kevin Michael
AU - Sadasivam, Rajani
AU - Chen, Ping
N1 - Publisher Copyright:
© 2017 IEEE.
PY - 2017/8/30
Y1 - 2017/8/30
N2 - The primary goal of this paper is to generate simulated data which is useful when only limited data is available. We introduce two techniques in this study to augment the datasets: (1) change a very small set of fields' values randomly and (2) through using generative adversarial networks (GANs). We propose a few analysis methods on classification problems to improve the accuracy of a well-sought class: (1) remove border samples between two categories, (2) reduce dimensionality through feature selection, (3) sacrifice the accuracy of less-valuable classes. We applied these methods to a real-word dataset: Smoking Cessation groups. One of the biggest challenges in this vein is that there is little available data, which is often the case in medical fields where data collection can be expensive and difficult. Also, a small amount of data may not contain sufficient information for machine learning methods to generate generalizable results. There are many existing methods to deal with this problem. However, their performance needs to be significantly improved in practice. Our results show that applying each of these analysis methods improves classification accuracy of the well-sought class and proved the GANs can generate many simulation data.
AB - The primary goal of this paper is to generate simulated data which is useful when only limited data is available. We introduce two techniques in this study to augment the datasets: (1) change a very small set of fields' values randomly and (2) through using generative adversarial networks (GANs). We propose a few analysis methods on classification problems to improve the accuracy of a well-sought class: (1) remove border samples between two categories, (2) reduce dimensionality through feature selection, (3) sacrifice the accuracy of less-valuable classes. We applied these methods to a real-word dataset: Smoking Cessation groups. One of the biggest challenges in this vein is that there is little available data, which is often the case in medical fields where data collection can be expensive and difficult. Also, a small amount of data may not contain sufficient information for machine learning methods to generate generalizable results. There are many existing methods to deal with this problem. However, their performance needs to be significantly improved in practice. Our results show that applying each of these analysis methods improves classification accuracy of the well-sought class and proved the GANs can generate many simulation data.
KW - classification
KW - data evaluation
KW - GANs
KW - Recruitment analysis
UR - https://www.scopus.com/pages/publications/85031744099
UR - https://www.scopus.com/pages/publications/85031744099#tab=citedBy
U2 - 10.1109/ICBK.2017.29
DO - 10.1109/ICBK.2017.29
M3 - Conference contribution
AN - SCOPUS:85031744099
T3 - Proceedings - 2017 IEEE International Conference on Big Knowledge, ICBK 2017
SP - 202
EP - 207
BT - Proceedings - 2017 IEEE International Conference on Big Knowledge, ICBK 2017
A2 - Lu, Ruqian
A2 - Wu, Xindong
A2 - Ozsu, Tamer
A2 - Wu, Xindong
A2 - Hendler, Jim
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 9 August 2017 through 10 August 2017
ER -