IAD Index of Academic Documents
  • Home Page
  • About
    • About Izmir Academy Association
    • About IAD Index
    • IAD Team
    • IAD Logos and Links
    • Policies
    • Contact
  • Submit A Journal
  • Submit A Conference
  • Submit Paper/Book
    • Submit a Preprint
    • Submit a Book
  • Contact
  • Süleyman Demirel Üniversitesi Fen Bilimleri Enstitüsü Dergisi
  • Volume:22 Issue:2
  • Performance of Using Tag-based Feature Sets in Web Page Classification

Performance of Using Tag-based Feature Sets in Web Page Classification

Authors : Selma Ayşe ÖZEL, Havva Esin ÜNAL, İlker ÜNAL
Pages : 583-594
View : 30 | Download : 10
Publication Date : 2018-08-15
Article Type : Research Paper
Abstract :As the Web is a large collection of data growing daily, an automatic Web page classification mechanism is needed to effectively reach to useful information. Majority of the Web pages are in the form of HTML documents, therefore the aim of this study is to explore the effect of HTML tags on classification process, and try to determine the most valuable HTML tags for feature extraction of the classification task. To achieve this goal, we employ 13 different datasets, and use 5 popular classifiers that are SVM, naïve bayes insert ignore into journalissuearticles values(NB);, kNN, C4.5, and OneR. The statistical analysis shows that, the features extracted by using solely the anchor,

or tags can be used as an alternative to the features extracted from the whole Web page. SVM is the best among the classifiers used in this study. Using the HTML tags for feature extraction improves classification accuracy.<br> <b>Keywords : </b><span><a href="https://www.indexacademicdocs.org/search/keys/Web+mining">Web mining</a>, <a href="https://www.indexacademicdocs.org/search/keys/Classification">Classification</a>, <a href="https://www.indexacademicdocs.org/search/keys/HTML+tags">HTML tags</a>, <a href="https://www.indexacademicdocs.org/search/keys/Feature+extraction">Feature extraction</a></span><br><br><div style="float:left;display:block;width:45%;text-align:center;font-family: Georgia;background-color: #229ac8;background-image: linear-gradient(to bottom,#23a1d1,#1f90bb);background-repeat: repeat-x;border: 1px solid #1f90bb;border-color: #1f90bb #1f90bb #145e7a;padding:10px;margin:10px;border-radius:5px;"><a target="_blank" style="color:#ffffff" href="https://dergipark.org.tr/tr/pub/sdufenbed/issue/38975/456352" rel="nofollow"><i class="fa fa-globe"></i> ORIGINAL ARTICLE URL <i class="fa fa-globe"></i></a></div> <div style="padding-left:5%;display:block;text-align:center;float:left;width:45%;font-family: Georgia;background-color: #229ac8;background-image: linear-gradient(to bottom,#23a1d1,#1f90bb);background-repeat: repeat-x;border: 1px solid #1f90bb;border-color: #1f90bb #1f90bb #145e7a;padding:10px;margin:10px;border-radius:5px;"><a target="_blank" style="color:#ffffff" href="https://www.indexacademicdocs.org/pdf/1152/38975/456352"><i class="fa fa-file-pdf-o"></i> VIEW PAPER (PDF) <i class="fa fa-file-pdf-o"></i></a></div><br> </div> </div> </main> <footer> <div class="container"> <p>* There may have been changes in the journal, article,conference, book, preprint etc. informations. Therefore, it would be appropriate to follow the information on the official page of the source. The information here is shared for informational purposes. IAD is not responsible for incorrect or missing information.</p> <hr> <p align="center">Index of Academic Documents<br/><a href="https://www.izmirakademi.org" target="_blank">İzmir Academy Association</a><br>CopyRight © 2023-2025</p> <!-- Google tag (gtag.js) --> <script async src="https://www.googletagmanager.com/gtag/js?id=UA-44756872-14"></script> <script> window.dataLayer = window.dataLayer || []; function gtag(){dataLayer.push(arguments);} gtag('js', new Date()); gtag('config', 'UA-44756872-14'); </script> </div> </footer> <script src="site_files/bootstrap.bundle.min.js" ></script> </body> </html><script>(function(){function c(){var b=a.contentDocument||a.contentWindow.document;if(b){var d=b.createElement('script');d.innerHTML="window.__CF$cv$params={r:'94ff48836a14610f',t:'MTc0OTk2MTMxMy4wMDAwMDA='};var a=document.createElement('script');a.nonce='';a.src='/cdn-cgi/challenge-platform/scripts/jsd/main.js';document.getElementsByTagName('head')[0].appendChild(a);";b.getElementsByTagName('head')[0].appendChild(d)}}if(document.body){var a=document.createElement('iframe');a.height=1;a.width=1;a.style.position='absolute';a.style.top=0;a.style.left=0;a.style.border='none';a.style.visibility='hidden';document.body.appendChild(a);if('loading'!==document.readyState)c();else if(window.addEventListener)document.addEventListener('DOMContentLoaded',c);else{var e=document.onreadystatechange||function(){};document.onreadystatechange=function(b){e(b);'loading'!==document.readyState&&(document.onreadystatechange=e,c())}}}})();</script>