· 8 years ago · Feb 28, 2018, 04:02 PM
1\documentclass[11pt]{article}
2
3\usepackage{graphicx}
4\usepackage[utf8]{inputenc}
5\usepackage[T1]{fontenc}
6
7\usepackage{mathpazo}
8
9
10\usepackage[a4paper, margin=1in]{geometry}
11
12\begin{document}
13\begin{titlepage}
14 \newcommand{\HRule}{\rule{\linewidth}{0.5mm}}
15 \center
16 \textsc{\LARGE University of Dhaka}\\[1.5cm]
17 \includegraphics[width=0.2\textwidth]{logo.png}\\[1cm]
18 \textsc{\LARGE Department of Computer Science \& Engineering}\\[1.5cm]
19 \HRule\\[0.4cm]
20
21 {\huge\bfseries Project Title: A Multimodal Approach for Fake News Detection}\\[0.4cm]
22 \HRule\\[1.5cm]
23 \begin{minipage}{0.4\textwidth}
24 \begin{flushleft}
25 \large
26 \textit{Author}\\
27 {Amar Debnath (Roll: 23)}\\
28 {Email:amar.csedu@gmail.com}
29 {Signature: }\\
30 \bigskip
31 \bigskip
32 {Redoan Rahman (Roll: 41)}\\
33 {Email:redoanrahman744@ gmail.com}\\
34 {Signature: }\\
35 \end{flushleft}
36 \end{minipage}
37 \begin{minipage}{0.4\textwidth}
38 \begin{flushright}
39 \begin{flushleft}
40 \large
41 \textit{Supervisor}\\
42 {Md. Mofijul Islam}\\
43 {Lecturer}\\
44 {Department of Computer Science and Engineering}\\
45 {University of Dhaka}\\
46 {Email: akash@cse.du.ac.bd}\\
47 {Signature: }\\
48 \bigskip
49 \bigskip
50 \end{flushleft}
51 \end{flushright}
52 \end{minipage}
53 \vfill\vfill\vfill
54 {\large\today}
55 \vfill
56
57\end{titlepage}
58
59\begin{samepage}
60\begin{Huge}
61 %\begin{center}
62 % \testbf{Project Proposal}
63 %\end{center}
64\end{Huge}
65
66\section*{Project Title: A Multimodal Approach for Fake News Detection }
67\section{Introduction}
68In this modern age of technology and Intelligent systems, humans have developed many techniques to cope with the problems of their surroundings. As many problems are being solved, more new problems are rising. Social media for news consumption is a double-edged sword. On the one hand, its low cost, easy access, and rapid dissemination of information lead people to seek out and consume news from social media. On the other hand, it enables the wide spread of “fake news", i.e., low quality news with intentionally false information. The extensive spread of fake news has the potential for extremely negative impacts on individuals and society. The extensive spread of fake news can have a serious negative impact on individuals and society.
69\\ First, fake news can break the authenticity balance of the news ecosystem. Second, fake news intentionally persuades consumers to accept biased or false beliefs. Fake news is usually manipulated by propagandists to convey political messages or influence. Third, fake news changes the way people interpret and respond to real news. For example, some fake news was just created to trigger people’s distrust and make them confused, impeding their abilities to differentiate what is true from what is not.
70\\To help mitigate the negative effects caused by fake news both to benefit the public and the news ecosystem It’s critical that we develop methods to automatically detect fake news on social media. During United States Election of 2016, there were 156 fake news that were identified later, among which 115 were pro-trump and 41 of them were pro-Clinton\cite{fake}.Therefore, fake news detection on social media has recently become an emerging research that is attracting tremendous attention.
71\subsection{Problem Definition}
72Several projects and researches are being conducted and it has grown to be a trending topic to detect fake news and reviews as well as trying to increase the accuracy of the existing models. Because of the possible mischief fake news causes to social media, It’s quite an urgent situation to deal with this issue in this modern age. Automatic fake news detection is a challenging problem in deception detection, and it has tremendous real-world political and social impacts. However, statistical approaches to combating fake news has been dramatically limited by the lack of labeled benchmark data sets.
73\subsection{Motivation}
74This era is driven by the power of Artificial Intelligence which is making the lives of humans more convenient. But it’s still a long way to go before our network can identify all the malicious or misleading news. If we can look around our society we can see that this false news is not only confusing people but also doing harm as well. For example: last year, a man carried an AR-15 rifle and walked in a Washington DC Pizzeria, because he recently read online that “this pizzeria was harboring young children as sex slaves as part of a child abuse ring led by Hillary Clintonâ€\cite{pizzaria}. The man was later arrested by police, and he was charged for firing an assault rifle in the restaurant. Besides some report also shows that Russia has created fake accounts and social bots to spread false stories\cite{TIME}. So we want to work on improving the fake news detection system.
75\bigskip
76\subsection{Objective}
77For designing an intelligent and accurate fake news detection system we must first find a moderate labeled data set to work on. We must experiment on that data set to evaluate our training model. After managing a labeled data set we will then develop and apply various machine learning algorithms to find a tolerable accuracy for our fake news detection system.
78\section{Related Works}
79Since it’s an emerging topic, There has been a lot of works related to fake review or fake news detection and judgment. Below we describe some of the papers that where many techniques and approaches are discussed.
80\subsection{Fake Review}
81Advertisers, marketers, and other stakeholders have motivation to produce fake positive user reviews for products they wish to promote or fake negative user reviews for products which they wish to disparage.In a fake user review, an actor will create a user account based on some marketing persona and post a user review purporting to be a real person with the traits of the persona.This is a misuse of the user review system, which universally only invite reviews from typical users and not paid fake personalities.
82\subsubsection{Behavioral Analysis of Review Fraud: Linking Malicious Crowd sourcing to Amazon and Beyond\cite{behaivour}}
83\par User reviews are a cornerstone of how we make decisions. From deciding what movies to view, products to purchase, restaurants to patronize, and even doctors to visit, user review aggregators like Amazon, Netflix, and Yelp shape our experiences. And yet, these reviews are vulnerable to manipulation. This manipulation threatens to degrade trust in these online platforms and in their products and services
84\paragraph{Proposed Solution}
85In this paper the authors focus on tasks posted to a single crowd sourcing site – Rapid Workers – that target Amazon. Typically, tasks on these sites pay workers from $0.10 to $1.50 per task, where a single target (e.g., a product on Amazon) may be subject to dozens of fake reviews launched from these crowd sourcing sites. The data set they made contains the following information: product ID, review ID, reviewer ID, review title, review content, rating, time-stamp and “verified purchase†flag. In total, they identified 5,200 unique reviewers and 350,000 unique reviews. They consider a reviewer to be a fraudulent reviewer if they have reviewed two or more products that have been targeted by a crowd sourcing effort and non-fraudulent otherwise. Then they made assumption that fraud reviewers tend to stay with their review style and make a lot of similar comment and word combinations. Keeping this assumption, they then find the Jaccard similarities and made proper calculations to justify their assumption.
86\paragraph{Future Works}
87This behavioral approach for the reviewers open a new path for the novel fake review detection architecture. This observation can be applied not only to online stores like amazon, but also various app stores and blogs as well. Besides, observing their linguistic evolution more carefully, we can find more insights into the behaviors of the fraud reviewers.
88\bigskip
89\subsubsection{Classification of Fake Product Ratings Using a Time line Based Approach\cite{timeline}}
90\par Detection of fake review and reviewers is currently a challenging problem in cyberspace. It is challenging primarily due to the dynamic nature of the methodology used to fake the review. There are several aspects to be considered when analyzing reviews to classify them effective into genuine and fake. Sentiment analysis, opinion mining and intend mining are fields of research that try to accomplish the goal through Natural Language Processing of the text content of the review.
91\paragraph{Proposed Solution}
92The three primary factors used in the proposed strategy to classify fake and genuine review
93ratings presented in this paper include: Time span between the first and last review, Average review rating over the time span, Number of review ratings. The Amazon Product Data set was used for the research analysis presented in this subsection. The data set was obtained from SNAP - Stanford Network Analysis Platform, a general purpose network analysis and graph mining library\cite{graph}.They then applied their custom algorithm which involves - Sorting of the time stamps in ascending order, Determination of the inter review period in Unix time by calculating the difference between the time periods of two consecutive ratings, Conversion of the time period of each rating into seconds, minutes and then days, thereby attaining the time interval between each rating in days, Calculation of the average of all the ratings for each time-period, Plotting of a graph with the information deduced, with average ratings for each of the time intervals and Calculation of the average of all ratings for a product. Using their algorithm they calculated the estimated average rating of the products.
94\paragraph{Future Works}
95With a growing trend in online shopping there is a need currently to provide an authentic environment for analyzing product reviews and ratings. This research paper expounds a methodology to identify fake ratings among genuine ones over time. This can be used to implicitly identify fake reviewers in cyberspace. A simple classification tool to identify the product specific point on the time line has also been used in the research work presented in this paper. The proposed fake rating filter therefore helps to tag possible fake reviewers, besides indicating the optimal period for the rating assessment for each product.
96\subsection{Fake News}
97\subsubsection{“Liar, Liar Pants on Fireâ€: A New Benchmark Data set for Fake News Detection\cite{LIAR}}
98\par Automatic fake news detection is a challenging problem in deception detection, and it has tremendous real-world political and social impacts. However, statistical approaches to combating fake news has been dramatically limited by the lack of labeled benchmark data sets. There are a lot of public data set of political or other news but most of them are not properly labeled or have much low content for the machine learning algorithms to trains.
99\paragraph{Proposed Solution}
100In this paper A new data set is proposed “LIAR†a new, publicly
101available data set for fake news detection. The authors collected a decade-long, 12.8K manually labeled short statements in various contexts from POLITIFACT.COM, which provides detailed analysis report and links to source documents for each case. This data set can be used for fact-checking research as well. Notably, this new data set is an order of magnitude larger than previously largest public fake news data sets of similar type. Empirically, they investigate automatic fake news detection based on surface-level linguistic patterns. They have designed a novel, hybrid convolutional neural network to integrate meta data with text. They showed that this hybrid approach can improve a text-only deep learning model.
102\bigskip
103\paragraph{Future Works}
104LIAR’s authentic, real-world short statements from various contexts with diverse speakers also make the research on developing broad-coverage fake news detector possible. When combining meta-data with text, significant improvements can be achieved for fine-grained fake news detection. Given the detailed analysis report and links to source documents in this data set, it is also possible to explore the task of automatic fact-checking over knowledge base in the future. The corpus can also be used for stance classification, argument mining, topic modeling, rumor detection, and political NLP research. Besides, The accuracy involving the features of the data set is not too high to apply in real life situations, so they data set can be modified or incorporated in a way such that the accuracy of the fake news detection can be improved.
105\subsubsection{"Fake News Detection Through Multi-Perspective Speaker Profiles"}
106\par The data set LIAR\cite{notunLIAR} given by professor William Yang Wang is an astounding empirical labeled data set which helps for the fake news detection research. But the features given in the data set is not enough to make a promising accuracy for detection. This paper proposes a novel method to incorporate speaker profiles into an attention based LSTM model for fake news detection. By adding more features, the performance metric of the learning model can be improved. Despite having a perfect labeled characteristic, the LIAR data set is still insufficient to help establish a fake news detection system which can be applied to real world.
107
108\paragraph{Proposed Solution}
109The solution of this paper involves incorporating speaker profiles into an LSTM model for fake news detection. Speaker profiles contribute to the model in
110two ways. One is to include them in the attention model. The other includes them as additional input data. By adding speaker profiles such as party affiliation, speaker title, location and credit history, the proposed model outperforms the state-of-the-art method by 14.5\% in accuracy using a benchmark fake news detection data set. This proves that speaker profiles provide valuable information to validate the credibility of news articles. evaluation is performed using the LIAR data set by Wang\cite{notunLIAR}. The maximum accuracy gained by LIAR data set is almost 27\%, and the maximum accuracy of the proposed LSTM model in this paper is around 41.5\%.
111
112\paragraph{Future Works}
113Augmenting speaker profiles have proven to improve the learning models, this not only adds extra accuracy but opens more possibilities that we can corporate other important characteristics of speakers and other attributes as well. Besides this hybrid LSTM model can be used in other applications like calculating the ratings of products and whether a person's speech can be trusted or not.
114\bigskip
115\section{Proposed Solution}
116 There is a decent data set consisting of a decade-long, 12.8K manually labeled short statements in various contexts from POLITIFACT.COM, which provides detailed analysis report and links to source documents for each case\cite{LIAR}. This data set can be used for fact-checking research as well. Notably, this new data set is an order of magnitude larger than previously largest public fake news data sets of similar type. But novel approaches in this data set results in quite a low accuracy which cannot be applied to real world situations. Despite using powerful learning algorithms, promising results still haven’t been found to fight with the fake news growth.
117 \par We propose to design a learning system producing results with acceptable accuracy so that it can be used in practical life. For this purpose, we aim to to design a multimodal data set by augmenting different new features into the already existing data sets. Since the existing data sets do not output good accuracy even when the best available learning algorithms are used, we think that a new improved data set is required to acquire better accuracy. In order to do that, we will need to augment different best-suited features into the available data sets to create a multimodal data set.
118 \par With the help of the multimodal data set we hope to design a good learning system that will perform with good accuracy. To do this, we will have to select features that are best-suited for the purpose from the data set. We will need to find the best combination of the available algorithms such as Support Vector Machine (SVM), Bi-directional Long Short Term Memory model( Bi-LSTM), Convolutional Neural Network (CNN), Recurrent Neural Network (RNN) etc. with the available set of features to be used.
119 \par After designing the learning system, we want to provide an interface to use it. For this we plan to create a website or a web application so that it is available to the general people.
120
121\end{samepage}
122\bigskip
123\begin{thebibliography}{9}
124\bibitem{fake}
125Journal of Economic Perspectives: \textit{Volume 31 Number 2—Spring 2017—Page 212}
126\bibitem{pizzaria}
127NYTIMES: In Washington Pizzeria Attack, Fake News Brought Real Guns
128\textit{https://www.nytimes.com/2016/12/05/business/media/comet-ping-pong-pizza-shooting-fake-news-consequences.html}. Accessed: 28-02-2018
129\bibitem{TIME}
130TIME: Inside Russia’s Social Media War on America
131\textit{http://time.com/4783932/inside-russia-social-media-waramerica/}.
132\bibitem{behaivour}. Accessed: 28-02-2018
133Kaghazgaran, Parisa, James Caverlee, and Majid Alfifi. "Behavioral Analysis of Review Fraud: Linking Malicious Crowdsourcing to Amazon and Beyond." \textit{ICWSM. 2017.}
134\bibitem{timeline}
135Thomas, Neha, and Susan Elias. "Classification of Fake Product Ratings Using a Timeline Based Approach." \textit{International Journal of Business Administration and Management Research 3.2 (2017): 12-15.}
136\bibitem{graph}
137Jure Leskovec and Andrej Krevl, “SNAP Datasets: Stanford Large Network Dataset Collectionâ€, June 2014
138\bibitem{LIAR}
139Wang, William Yang. "" Liar, Liar Pants on Fire": A New Benchmark Dataset for Fake News Detection." \textit{arXiv preprint arXiv:1705.00648 (2017)}.
140\bibitem{notunLIAR}
141Long Y, Lu Q, Xiang R, Li M, Huang CR. "Fake News Detection Through Multi-Perspective Speaker Profiles". In Proceedings of the Eighth International Joint Conference on Natural Language Processing (Volume 2: Short Papers) 2017 (Vol. 2, pp. 252-256). Accessed: February 2018
142\end{thebibliography}
143\end{document}