<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.0 20120330//EN" "JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta><journal-id journal-id-type="other">Journal</journal-id>
      <journal-title-group>
        <journal-title>International Journal of Scientific Research and Management, IJSRM</journal-title>
      </journal-title-group>
    <publisher><publisher-name>Academic Publisher</publisher-name></publisher></journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.18535/ijsrm/v7i1.ec01</article-id>
      <title-group>
        <article-title>A data-oriented approach for outlier detection</article-title>
      </title-group>
      <contrib-group content-type="author">
        <contrib contrib-type="author">
          <name><given-names>Nripesh Trivedi</given-names></name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        
      </contrib-group>
      <aff id="aff0">
          
          <institution content-type="orgname">Department of mathematical sciences. Indian Institute of Technology</institution>
          ,
          <addr-line>Varanasi</addr-line>
          ,
          <country country="IN">India</country>
        </aff><pub-date>
        <year>2000</year>
      </pub-date>
      <volume>07</volume>
      <issue>01</issue>
      <fpage>2</fpage>
      <lpage>4</lpage>
      <abstract>
        <p>In this paper, characteristics of data obtained from the sensors (used in OpenSense project) are identified in order to build a data-oriented approach. This approach consists of application of Class Outliers: Distance Based (CODB) and Hoeffding tree algorithms. Subsequently, machine learning models were built to detect outliers in a sensor data stream. The approach presented in this paper may be used for developing methodologies for data-oriented outlier detection.</p>
      </abstract>
      <kwd-group>
        <kwd>Data characteristics</kwd>
        <kwd>Nominal attribute</kwd>
        <kwd>outlier analysis</kwd>
        <kwd>machine learning</kwd>
        <kwd>model verification</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
          <sec>
        <title>
        
          Introduction
        </title>
        <p>Numerous algorithms have been proposed to detect outliers. However, for this paper, the approach adopted is exclusive to the data-set under consideration. Since the approach is meant exclusively for the data-set, it may be termed as a data-oriented approach. This paper may be the first to use to a data-oriented approach for outlier analysis.</p>
      </sec>
      <sec>
        <title>Outlier detection: a data-oriented methodology</title>
        
          <p>Data-set Description</p>
        
        <p>The Data-set used in this paper was made available by Dr. Jean Paul Calbimonte who is an active researcher in the OpenSense project. Size of the data-set is more than 16 million rows. Around 100000 rows of this data-set were used to build machine learning models while other rows of this data-set were used to build a data stream that these machine learning models could use in prediction. </p>
        <p>The attributes in the data-set are:</p>
        <list list-type="simple">
          <list-item>
            <p>Latitude</p>
          </list-item>
          <list-item>
            <p>Longitude</p>
          </list-item>
          <list-item>
            <p>LDSA</p>
          </list-item>
          <list-item>
            <p>Station</p>
          </list-item>
        </list>
        <p>In OpenSense project, sensors were positioned on top of the buses (public transportation) to measure pollution levels in Lausanne city. The term LDSA stands for lung deposited surface area. It is a way to measure the quantity of particles</p>
        <p>Table I: Data Attributes and their respective range of values</p>
        <table-wrap specific-use="rules" position="float" orientation="portrait">
          <table>
            <tr>
              <td rowspan="1" colspan="1">Latitude</td>
              <td rowspan="1" colspan="1">46.5202347 - 46.5218066</td>
            </tr>
            <tr>
              <td rowspan="1" colspan="1">Longitude</td>
              <td rowspan="1" colspan="1"> 6.6307456 - 6.6315791</td>
            </tr>
            <tr>
              <td rowspan="1" colspan="1">Station</td>
              <td rowspan="1" colspan="1">41, 43, 45, 47. 48, 49, 50, 51, 54, 55</td>
            </tr>
            <tr>
              <td rowspan="1" colspan="1">LDSA</td>
              <td rowspan="1" colspan="1">1 - 2000</td>
            </tr>
          </table>
        </table-wrap>
        <p>10 stations indicated in the table I include both static and mobile stations.</p>
      </sec>
      <sec>
        <title>Methodology for outlier detection</title>
        <p>Uniform Random sampling is a method for summarizing multidimensional data streams [<xref rid="R5" ref-type="bibr" id="ID981d55a3-f344-40f9-92b5-1ea977dd7ff8">1</xref>]. Using this method, data-sets could be sampled. This sampling was applied on the data-set described in Table I. The data-set obtained from sampling (say data-set I) was treated in the following manner: </p>
        <p>CODB Algorithm was applied over data-set I to detect outliers. Further, data-set I and outliers detected in this data were used for building machine learning models that could learn from a data stream and find outliers in it.</p>
        <p>Station is a numerical attribute but it may be used as a class (nominal) attribute. The advantage of using Station attribute as a class attribute is that class labels may be assigned to each row in the original data-set and also data-set I (the data-set obtained from sampling). An instance in Data-set I is of the form (Station, Latitude, Longitude, LDSA) where Station is a nominal attribute and other attributes are numerical.</p>
        <p>Class Outliers: Distance Based (CODB) Algorithm was applied over the data-set I in the following manner: </p>
        <list list-type="simple">
          <list-item>
            <p>Since Station attribute is a nominal attribute; Latitude, Longitude and LDSA attributes are independent of the corresponding Station data value (values are in Table I). Since data values of station attribute are independent of each other as station is a nominal attribute, data values of Latitude, Longitude and LDSA attributes for each station are also independent of the corresponding station data values. Each Station value is also independent of all other station value (as stated above).  Thus, when CODB algorithm is applied to data-set I, detection of outliers is independent of the station value. </p>
          </list-item>
          <list-item>
            <p>Data values from to one station were paired with another station since LDSA, Latitude and Longitude attributes have almost similar numerical ranges for both of these stations. Station values paired together are shown in table II while corresponding LDSA, Latitude and Longitude values are shown in figure 1 and Table I respectively. Due to this pairing, CODB algorithm does not need to be applied individually to data values from each station. It could be individually applied to five pairs of stations Moreover, since decision to use class attribute as a nominal attribute is crucial to the application of CODB and also of Hoeffding tree ((VFDT) later in the paper), the data is involved in a central role in the analysis.</p>
          </list-item>
        </list>
        
          <p> Table II: Paired Stations in Data-set I</p>
        <table-wrap specific-use="rules" position="float" orientation="portrait">
          <table>
            <tr>
              <td rowspan="1" colspan="1">Pair number</td>
              <td rowspan="1" colspan="1"> Station number 1</td>
              <td rowspan="1" colspan="1"> Station number 2</td>
            </tr>
            <tr>
              <td rowspan="1" colspan="1"> 1</td>
              <td rowspan="1" colspan="1">49</td>
              <td rowspan="1" colspan="1">50</td>
            </tr>
            <tr>
              <td rowspan="1" colspan="1">2</td>
              <td rowspan="1" colspan="1">45</td>
              <td rowspan="1" colspan="1">54</td>
            </tr>
            <tr>
              <td rowspan="1" colspan="1">3</td>
              <td rowspan="1" colspan="1">41</td>
              <td rowspan="1" colspan="1">47</td>
            </tr>
            <tr>
              <td rowspan="1" colspan="1">4</td>
              <td rowspan="1" colspan="1">43</td>
              <td rowspan="1" colspan="1">55</td>
            </tr>
            <tr>
              <td rowspan="1" colspan="1">5</td>
              <td rowspan="1" colspan="1">48</td>
              <td rowspan="1" colspan="1"> 51</td>
            </tr>
          </table>
        </table-wrap>

        
                  
<fig id="fig1">
<label>Figure 1:</label>
<caption>
<title>
Distribution of LDSA attributes for pairs of stations for data-set I</title>
</caption>
<graphic xlink:href="fig1.png"/>
</fig>


        <list list-type="simple">
          <list-item>
            <p>CODB algorithm has three components: </p>
          </list-item>
          <list-item>
            <p>PCL (T, K)</p>
          </list-item>
          <list-item>
            <p>Dev (T) </p>
          </list-item>
          <list-item>
            <p>K-Dist (T)</p>
          </list-item>
        </list>
        
          <p>Where T is the instance for which COF (T) (degree of being a class outlier) is evaluated and K is number of nearest neighbors. While evaluating Dev (T) and K-Dist (T), three numerical attributes within the data-set are used, namely, Latitude, Longitude and LDSA. The value of PCL (T, K) is a number between 0 and K and indicates nearest neighbors that belong to the same class (station attribute). CODB algorithm was applied over data values belonging to each of the five pairs of Stations in the manner described above. After running initial experiments on data-set I, maximum value of deviation (Dev (T))) was found to lie in thousands and maximum value of K-Dist (K-Dist (T)) was 0.0. Since maximum value of K-Dist was 0.0, corresponding parameter may be any arbitrary real number between 0 and 1 as suggested in [<xref rid="R6" ref-type="bibr" id="IDf9d327a6-4b7a-48c6-b8f2-f17631832b9f">2</xref>]. Number of nearest neighbors (K) were set to 7 as suggested in [<xref rid="R6" ref-type="bibr" id="IDf3780bf1-7b02-4122-beb6-30523c1f7e7a">2</xref>]. The parameters with their corresponding values are shown in the table III below. Given the varying quality of measurements obtained from different sensors used in the OpenSense project, 10% of the data values in data-set were regarded as outliers.</p>
        
        <p>Table III: Parameters with their respective values</p>
        <table-wrap specific-use="rules" position="float" orientation="portrait">
          <table>
            <tr>
              <td rowspan="1" colspan="1">Parameters</td>
              <td rowspan="1" colspan="1">Values</td>
            </tr>
            <tr>
              <td rowspan="1" colspan="1"> </td>
              <td rowspan="1" colspan="1"> 1000</td>
            </tr>
            <tr>
              <td rowspan="1" colspan="1"> </td>
              <td rowspan="1" colspan="1">0.1</td>
            </tr>
            <tr>
              <td rowspan="1" colspan="1"> k</td>
              <td rowspan="1" colspan="1"> 7</td>
            </tr>
          </table>
        </table-wrap>
        <p>IV Outlier Detection In Data Stream</p>
        <p>For detecting outliers in the data stream (detailed description of the data stream may be found in data-set description), an algorithm should be chosen that should build machine learning models using a small sample of data and these machine learning models may be used for prediction. Further, the machine learning models should learn while carrying out prediction. Hoeffding tree (VFDT) provides a robust solution [<xref rid="R7" ref-type="bibr" id="ID39fcfa38-bd34-46e6-86ba-2a3b14c1af85">3</xref>] for this requirement. Hoeffding tree was trained over data-set I and also over the outliers detected in data-set I. A separate Hoeffding tree algorithm was trained over five pair of stations shown in table (IV), therefore, five machine learning model were obtained. The efficiency of each model is shown in the table below.</p>
        <p>Table IV: Efficiency for pairs of station</p>
        <table-wrap specific-use="rules" position="float" orientation="portrait">
          <table>
            <tr>
              <td rowspan="1" colspan="1">Pair number</td>
              <td rowspan="1" colspan="1">Overall Efficiency</td>
            </tr>
            <tr>
              <td rowspan="1" colspan="1"> 1</td>
              <td rowspan="1" colspan="1"> 92.66%</td>
            </tr>
            <tr>
              <td rowspan="1" colspan="1"> 2</td>
              <td rowspan="1" colspan="1"> 96.21%</td>
            </tr>
            <tr>
              <td rowspan="1" colspan="1">3</td>
              <td rowspan="1" colspan="1"> 95.32%</td>
            </tr>
            <tr>
              <td rowspan="1" colspan="1"> 4</td>
              <td rowspan="1" colspan="1"> 95.30%</td>
            </tr>
            <tr>
              <td rowspan="1" colspan="1">5</td>
              <td rowspan="1" colspan="1">97.73%</td>
            </tr>
          </table>
        </table-wrap>
        <p>Since minority class (outlier class) is 10% of data-set I while majority class (i.e. non- outlier class) is 90% of data-set I, precision and recall must be used to verify the accuracy of the five machine learning models. In order to do so, 10 fold cross validation is applied. Table (V) shows the precision and recall for each pair of stations. </p>
        <p>Table V: Precision and Recall value for each pair</p>
        <table-wrap specific-use="rules" position="float" orientation="portrait">
          <table>
            <tr>
              <td rowspan="1" colspan="1">Pair number</td>
              <td rowspan="1" colspan="1"> Precision</td>
              <td rowspan="1" colspan="1"> Recall</td>
            </tr>
            <tr>
              <td rowspan="1" colspan="1"> 1</td>
              <td rowspan="1" colspan="1">87.38%</td>
              <td rowspan="1" colspan="1">93.26%</td>
            </tr>
            <tr>
              <td rowspan="1" colspan="1">2</td>
              <td rowspan="1" colspan="1">81.85%</td>
              <td rowspan="1" colspan="1">93.97%</td>
            </tr>
            <tr>
              <td rowspan="1" colspan="1">3</td>
              <td rowspan="1" colspan="1">83.90%</td>
              <td rowspan="1" colspan="1">91.92%</td>
            </tr>
            <tr>
              <td rowspan="1" colspan="1">4</td>
              <td rowspan="1" colspan="1">92.66%</td>
              <td rowspan="1" colspan="1">95.81%</td>
            </tr>
            <tr>
              <td rowspan="1" colspan="1">5</td>
              <td rowspan="1" colspan="1">84.35%</td>
              <td rowspan="1" colspan="1"> 98.36%</td>
            </tr>
          </table>
        </table-wrap>
      </sec>
      <sec>
        <title>Discussion</title>
        <p>This discussion is about the generalizability of the approach. Since, Uniform random sampling was applied to the-data-set, the complete data-set need not be inspected for discovering patterns in the data-set. As described in the methodology for outlier detection section, each row in data-set is independent of all other rows due to treatment of station attribute. Therefore, relation among attributes of the data-set are identified as shown in methodology for outlier detection section, thus making the approach data-oriented. This approach is contrary to the approach adopted by the authors in [<xref rid="R8" ref-type="bibr" id="ID05be7af3-3fbc-4782-91f6-a08d7df23e51">4</xref>].The authors adopt an algorithmic approach where a novel algorithm, kernel k-means is used for unsupervised clustering rather than the traditional k-means method. </p>
      </sec>
      <sec>
        <title>Conclusion and Future work</title>
        <p>The approach shown in this paper is specific to OpenSense data. In order to come up with generalized approaches, it is necessary to find common properties in data obtained from various sensor networks. These common properties could be used to design generalized approaches. </p>
      </sec>
  </body>
  <back>
    <ref-list>

<title>References</title>
<ref id="R5"><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Appice</surname><given-names>Annalisa</given-names></name><name><surname>Ciampi</surname><given-names>Anna</given-names></name><name><surname>Fumarola</surname><given-names>Fabio</given-names></name><name><surname>Malerba</surname><given-names>Donato</given-names></name></person-group><article-title>Data Mining Techniques in Sensor Networks</article-title><year>2014</year><pub-id pub-id-type="doi">10.1007/978-1-4471-5454-9</pub-id></element-citation></ref><ref id="R6"><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>.</surname><given-names>NabilMHewahi</given-names></name></person-group><article-title>Intelligent Tutoring System: Hierarchical Rule as a Knowledge Representation and Adaptive Pedagogical Model</article-title><source>Information Technology Journal</source><year>2007-may</year><fpage>739</fpage><lpage>744</lpage><pub-id pub-id-type="doi">10.3923/itj.2007.739.744</pub-id></element-citation></ref><ref id="R7"><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Hulten</surname><given-names>Geoff</given-names></name><name><surname>Spencer</surname><given-names>Laurie</given-names></name><name><surname>Domingos</surname><given-names>Pedro</given-names></name></person-group><article-title>Mining time-changing data streams</article-title><source>Proceedings of the seventh ACM SIGKDD international conference on Knowledge discovery and data mining - KDD \textquotesingle01</source><publisher-name>ACM Press</publisher-name><year>2001</year><pub-id pub-id-type="doi">10.1145/502512.502529</pub-id></element-citation></ref><ref id="R8"><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Trivedi</surname><given-names>Nripesh</given-names></name><name><surname>Asamoah</surname><given-names>DanielAdomako</given-names></name><name><surname>Doran</surname><given-names>Derek</given-names></name></person-group><article-title>Keep the conversations going: engagement-based customer segmentation on online social service platforms</article-title><source>Information Systems Frontiers</source><year>2016-nov</year><fpage>239</fpage><lpage>257</lpage><pub-id pub-id-type="doi">10.1007/s10796-016-9719-x</pub-id></element-citation></ref></ref-list>
  </back>
</article>
