Data Modeling

Mathematical and Computational Data Modeling | Online ISSN 3143-9217
2
Citations
10.2k
Views
45
Articles
Your new experience awaits. Try the new design now and help us make it even better
Switch to the new experience
RESEARCH ARTICLE   (Open Access)

BeamNet: A DOM-Based Approach to Personalized Web Content Curation and Real-Time Notification Delivery

Abstract 1. Introduction 2. Methodology 3. Results and Discussion 4. Conclusion References

Md. Mubdiur Rahman 1*

 

+ Author Affiliations

Data Modeling 2 (1) 1-8 https://doi.org/10.25163/data.2110896

Submitted: 23 May 2021 Revised: 09 July 2021  Accepted: 18 July 2021  Published: 20 July 2021 


Abstract

The volume of information generated across the web has grown to a point were locating what actually matters to an individual reader has become, somewhat paradoxically, harder rather than easier. This study presents BeamNet, a mobile-first system designed to let users assemble their own curated web feed and receive push notifications when the content they care about changes. prior approaches to web content extraction have relied variously on Document Object Model (DOM) tree traversal, text-pattern matching, and headless-browser rendering, each carrying its own trade-offs between accuracy and computational cost. BeamNet allows a user to select, from within a mobile application, multiple visible text elements on a target webpage; the Lowest Common Ancestor (LCA) of the selected elements is then computed and stored, together with the source URL, as a reusable extraction pattern. A server-side headless browser subsequently revisits each stored source at fixed intervals, re-applies the LCA pattern to retrieve updated content, and compares the result against previously stored data to detect change, triggering a Firebase Cloud Messaging notification when new content is found. Informal testing across a heterogeneous sample of websites indicated a content-extraction error rate of approximately 2%–5%, arising primarily from cases in which unrelated text elements shared the same LCA as the intended target. the findings suggest that a lightweight, LCA-based extraction strategy can support personalized, user-controlled content curation without requiring site-specific scraping rules, though scalability and extraction precision remain open engineering challenges that warrant further, more rigorously controlled evaluation.

Keywords: personalized web feed; web scraping; DOM-based content extraction; push notification; content curation

References

Bahuleyan, H., & Vechtomova, O. (2017). UWaterloo at SemEval-2017 Task 8: Detecting stance towards rumours with topic-independent features. https://github.com/HareeshBahuleyan/

Bulus, H. N., Uzun, E., & Doruk, A. (2017). Comparison of string matching algorithms in Web documents. In Proceedings of the International Scientific Conference (UNITECH) (Vol. 2, pp. 279–282).

Chakrabarti, S., Punera, K., & Subramanyam, M. (2002). Accelerated focused crawling through online relevance feedback. In Proceedings of the 11th International Conference on World Wide Web (WWW) (pp. 148–159).

Chakrabarti, S., van den Berg, M., & Dom, B. (1999). Focused crawling: A new approach to topic-specific Web resource discovery. Computer Networks, 31(11–16), 1623–1640.

Diligenti, M., Coetzee, F. M., Lawrence, S., Giles, C. L., & Gori, M. (2000). Focused crawling using context graphs. In Proceedings of the 26th International Conference on Very Large Data Bases (VLDB) (pp. 527–534).

Dong, H., & Hussain, F. K. (2013). SOF: A semi-supervised ontology-learning-based focused crawler. Concurrency and Computation: Practice and Experience, 25(12), 1755–1770.

Ebadi, N., Jozani, M., Choo, K. K. R., & Rad, P. (2022). A memory network information retrieval model for identification of news misinformation. IEEE Transactions on Big Data, 8(5), 1358–1370. https://doi.org/10.1109/TBDATA.2020.3048961

Ferrara, E., De Meo, P., Fiumara, G., & Baumgartner, R. (2014). Web data extraction, applications and techniques: A survey. Knowledge-Based Systems, 70, 301–323. https://doi.org/10.1016/j.knosys.2014.07.007

Gheorghe, M., Mihai, F.-C., & Dârdala, M. (2018). Modern techniques of web scraping for data scientists. Revista Româna de Interac?iune Om-Calculator, 11(1), 45–58.

Grubbs, F. E. (1969). Procedures for detecting outlying observations in samples. Technometrics, 11(1), 1–21.

Ito, T., Kim, Y., & Fukuta, N. (2015). The design and implementation of a personalized news recommendation system. IEEE Computer Society.

Najafabadi, M. M., Villanustre, F., Khoshgoftaar, T. M., Seliya, N., Wald, R., & Muharemagic, E. (2015). Deep learning applications and challenges in big data analytics. Journal of Big Data, 2(1), Article 1.

Najafabadi, M. M., Villanustre, F., Khoshgoftaar, T. M., Seliya, N., Wald, R., & Muharemagic, E. (2015). Deep learning applications and challenges in big data analytics. Journal of Big Data, 2(1), 1.

Nusret, H., Namik, B., Üniversitesi, K., Onyedi, B., & Üniversitesi, E. (2017). Comparison of string matching algorithms in web documents. ResearchGate. https://www.researchgate.net/publication/321228533

Pinkerton, B. (1994). Finding what people want: Experiences with the WebCrawler. In Proceedings of the 2nd International World Wide Web Conference (Vol. 94, pp. 17–20).

Salem, H., & Mazzara, M. (2020). Pattern matching-based scraping of news websites. Journal of Physics: Conference Series, 1694(1), Article 012011. https://doi.org/10.1088/1742-6596/1694/1/012011

Shchekotykhin, K., Jannach, D., & Friedrich, G. (2010). XCrawl: A high-recall crawling method for Web mining. Knowledge and Information Systems, 25(2), 303–326.

Uçar, E., Uzun, E., & Tüfekci, P. (2017). A novel algorithm for extracting the user reviews from Web pages. Journal of Information Science, 43(5), 696–712.

Uzun, E. (2020). A novel web scraping approach using the additional information obtained from web pages. IEEE Access, 8, 61726–61740. https://doi.org/10.1109/ACCESS.2020.2984503

Uzun, E., Agun, H. V., & Yerlikaya, T. (2013). A hybrid approach for extracting informative content from Web pages. Information Processing & Management, 49(4), 928–944.

Uzun, E., Bulus, H. N., Doruk, A., & Özhan, E. (2017). Evaluation of Hap, AngleSharp and HtmlDocument in Web content extraction. In Proceedings of the International Scientific Conference (UNITECH) (Vol. 2, pp. 275–278).

Uzun, E., Güner, E. S., Kiliçaslan, Y., Yerlikaya, T., & Agun, H. V. (2014). An effective and efficient Web content extractor for optimizing the crawling process. Software: Practice and Experience, 44(10), 1181–1199.

Wu, Y.-C. (2016). Language independent web news extraction system based on text detection framework. Information Sciences, 342, 132–149.

Wu, Y.-C. (2016). Language independent Web news extraction system based on text detection framework. Information Sciences, 342, 132–149.


Article metrics
View details
0
Downloads
0
Citations
6
Views

View Dimensions


View Plumx


View Altmetric



0
Save
0
Citation
6
View
0
Share