Posts

Showing posts with the label Big Data

Apache Hadoop | Implementation

Image
Starting with Apache Hadoop System analysis and Design This article explains how the system is analyzed to carry out the work for the proposed system. System analysis is the process of gathering and interpreting facts, diagnosing problems, and using the facts to improve the system. System analysis does more than just solve the current problem especially when there is no such system exists that is going to be developed. The future needs of the business and the changes required to meet the needs are analyzed. Once the decision is made, the plan is developed to implement the recommendations. The plan includes all system design features, such as new data capture needs(storage system), operating systems, equipment, and personal needs. The system design is like a blueprint: it specifies all the features that are to be in the finished product. Class diagram The class diagram is static. It represents the static view of an application. It...

Starting with Apache Hadoop

Image
Starting with Apache Hadoop In Hadoop, a single master is managing many slaves The master node consists of a JobTracker , Tasktracker , NameNode , and DataNode . A slave or worker node acts as both DataNode and TaskTracker though it is possible to have data-only worker node, and compute-only workerNodes. NameNode holds the file system metadata. The files are broken up and spread over the DataNode and JobTracker schedules and the manager's job. The TaskTracker executes the individual map and reduced function. If a machine fails, Hadoop continues to operate the cluster by shifting work to the remaining machines. The input file, which resides on a distributed file system throughout the cluster, is split into even-sized chunks replicated for fault tolerance. Haddopp divides each map to reduce jobs into a set of tasks. Each chunk of input is processed by a map task, which outputs a list of key-value pairs. In Hadoop, the shuffle phase o...

Handling of Big Data on the Internet

Image
Handling of Big Data on the Internet When we handle text files that are about 4000-5000 lines long, it is difficult, time-consuming, Input/Output overhead, non-scalable, hardware fault, unnecessary repetition of code, loss of memory space, difficult to process errors and prevent error propagation, etc. The term "Big Data" applies to the above data. Suppose I have a file that is in the English language and is about 4000-5000 lines long with an average of 100 characters per line with a few anomalies as blank lines Heuristics The heuristics we are going to sue are as follows: Articles (a, an, the) Prepositions (of, at, about, around, besides, aside, above, over) Conjunctions (and, between, or, because, hence, since, although, though, not only, but also, but, so, therefore) Adverbs ( adverbs are easy to recognize as they mostly have "ly" as their suffix Pronouns (I, he, we, our, their, he, she, it) ...