You are on page 1of 32

HBase – In Detail

Pushpinder Singh
PAXCEL Technologies
Hbase: Introduction
• Distributed
• Column-oriented
• Database
• Built on top of HDFS
Hbase: Introduction
• Is a hadoop application
• To be used when we require
– Real-time read/ write
– random-access to
– very large data sets
• Built to scale linearly just by adding nodes
Hbase: Introduction
• Is not relational
• Does not support SQL
• Given the proper problem space it is able to
do what an RDBMS cannot:
– Ably host very large & sparsely populated tables
on clusters made from commodity hardware
Typical RDBMS Scaling Story
• Initial public launch
– Move from local workstation to shared, remote hosted MySQL instance with a well-defined schema.
• Service becomes more popular; too many reads hitting the database
– Add memcached to cache common queries. Reads are now no longer strictly ACID; cached data must expire.
• Service continues to grow in popularity; too many writes hitting the database
– Scale MySQL vertically by buying a beefed up server with 16 cores, 128 GB of RAM, and banks of 15 k RPM hard
drives. Costly.
• New features increases query complexity; now we have too many joins
– Denormalize your data to reduce joins. (That’s not what they taught me in DBA school!)
• Rising popularity swamps the server; things are too slow
– Stop doing any server-side computations.
• Some queries are still too slow
– Periodically prematerialize the most complex queries, try to stop joining in most cases.
• Reads are OK, but writes are getting slower and slower
– Drop secondary indexes and triggers (no indexes?).
• More Options
– attempt to build some type of partitioning on your largest tables, or look into some of the commercial solutions
that provide multiple master capabilities.
What we end up with…
• something that is no longer a true RDBMS, sacrificing features and
conveniences for compromises and complexities.
• Any form of slave replication or external caching introduces weak
consistency into your now denormalized data.
• The inefficiency of joins and secondary indexes means almost all
queries become primary key lookups.
• A multiwriter setup likely means no real joins at all and distributed
transactions are a nightmare.
• There’s now an incredibly complex network topology to manage
with an entirely separate cluster for caching.
• Even with this system and the compromises made, you will still
worry about your primary master crashing and the daunting
possibility of having 10 times the data and 10 times the load in a
few months.
Hbase and its Scaling…
• No real indexes
– Rows are stored sequentially, as are the columns within each row. Therefore, no issues with
index bloat, and insert performance is independent of table size.
• Automatic partitioning
– As your tables grow, they will automatically be split into regions and distributed across all
available nodes.
• Scale linearly and automatically with new nodes
– Add a node, point it to the existing cluster, and run the regionserver. Regions will
automatically rebalance and load will spread evenly.
• Commodity hardware
– Clusters are built on $1,000–$5,000 nodes rather than $50,000 nodes. RDBMS are hungry I/O,
which is the most costly type of hardware.
• Fault tolerance
– Lots of nodes means each is relatively insignificant. No need to worry about individual node
downtime.
• Batch processing
– MapReduce integration allows fully parallel, distributed jobs against your data with locality
awareness
Hbase: A Map
• At its core, Hbase Table is a map
• Depending on your programming language background, you may be more
familiar with the terms
– associative array (PHP),
– dictionary (Python),
– Hash (Ruby), or
– Object (JavaScript).
• A map is "an abstract data type composed of a collection of keys and a
collection of values, where each key is associated with one value."
• Using JavaScript Object Notation, here's an example of a simple map
where all the values are just strings:
{ "zzzzz" : "woot",
"xyz" : "hello",
"aaaab" : "world",
"1" : "x",
"aaaaa" : "y" }
Hbase: sorted
• Unlike most map implementations, in HBase/BigTable the
key/value pairs are kept in strict alphabetical order.
• That is to say that the row for the key "aaaaa" should be right next
to the row with key "aaaab" and very far from the row with key
"zzzzz".
{ "1" : "x",
"aaaaa" : "y",
"aaaab" : "world",
"xyz" : "hello",
"zzzzz" : "woot" }
• Because these systems tend to be so huge and distributed, this
sorting feature is actually very important. The spacial propinquity of
rows with like keys ensures that when you must scan the table, the
items of greatest interest to you are near each other.
• Another example: domain name storage
Hbase: MultiDimentional
• Adding one dimension to our example gives us this:
{
"1" : {
"A" : "x"
},
"aaaaa" : {
"A" : "y"
},
"aaaab" : {
"A" : "world"
},
"xyz" : {
"A" : "hello”
},
"zzzzz" : {
"A" : "woot"
}
}
Hbase: MultiDimentional
Lets add another dimension
{ // ...
"aaaaa" : {
"A" : {
"foo" : "y",
"bar" : "d"
}
, "B" : {
"" : "w"
}
}
, "aaaab" : {
"A" : {
"foo" : "world",
"bar" : "domination"
}
, "B" : {
"" : "ocean" }
}
, // ... }
Hbase: Sparse
• A given row can have any number of columns
in each column family, or none at all.
• The other type of sparseness is row-based
gaps, which merely means that there may be
gaps between keys.
Hbase Data Model: First Look
• Applications store data into labeled tables.
• Tables are made of rows and columns.
• Table cells—the intersection of row and
column coordinates—are versioned.
• By default, their version is a timestamp auto-
assigned by HBase at the time of cell insertion.
• A cell’s content is an uninterpreted array of
bytes.
Hbase Data Model: First Look
• Table row keys are also byte arrays, so
theoretically anything can serve as a row key:
– from strings to binary representations of longs or
even serialized data structures.
• Table rows are sorted by row key, the table’s
primary key.
• By default, the sort is byte-ordered.
• All table accesses are via the table primary
key.
Hbase Data Model: First Look
• Row columns are grouped into column
families.
• All column family members have a common
prefix, so, for example, the columns
temperature:air and Temperature:dew_point
are both members of the temperature column
family, whereas station:identifier belongs to
the station family.
Hbase Data Model: First Look
• A table’s column families must be specified up
front as part of the table schema definition
• New column family members can be added on
demand. For example, a new column
station:address can be offered by a client as
part of an update, and its value persisted, as
long as the column family station is already in
existence on the targeted table.
Hbase Data Model: First Look
• Physically, all column family members are stored
together on the filesystem.
• Though earlier we described HBase as a column-
oriented store, it would be more accurate if it
were described as a column-family-oriented
store.
• Because tunings and storage specifications are
done at the column family level, it is advised that
all column family members have the same
general access pattern and size characteristics.
More About Hbase Data Model
• Regions
– Tables are automatically partitioned horizontally by HBase into regions.
– Each region comprises a subset of a table’s rows.
– A region is defined by its first row, inclusive, and last row, exclusive, plus a randomly generated
region identifier.
– Initially a table comprises a single region but as the size of the region grows, after it crosses a
configurable size threshold, it splits at a row boundary into two new regions of approximately
equal size.
– Until this first split happens, all loading will be against the single server hosting the original region.
– As the table grows, the number of its regions grows.
– Regions are the units that get distributed over an HBase cluster.
– In this way, a table that is too big for any one server can be carried by a cluster of servers with
each node hosting a subset of the table’s total regions.
– This is also the means by which the loading on a table gets distributed.
• Locking
– Row updates are atomic, no matter how many row columns constitute the row-level transaction.
– This keeps the locking model simple.
Hbase Implementation
• HBase setup is characterized with an HBase master
node orchestrating a cluster of one or more
regionserver slaves
Hbase Master Resposibility
• Boot-strap a new installation
• Assign regions to registered region servers
• Recover region server failures
Hbase Region Servers
• The regionservers carry zero or more regions
• Feed client read/write requests.
• They also manage region splits informing the
HBase master about the new daughter
regions.
• Regionserver slave nodes are listed in the
<$Hbase_Install>/conf/regionservers
configuration file
More about HBase
• Cluster site-specific configuration is made in the
<$Hbase_Install>/conf/hbase-site.xml and
<$Hbase_Install>/conf/hbase-env.sh files
• HBase persists data via the Hadoop filesystem API.
• Since there are multiple implementations of the
filesystem interface—
• one for the local filesystem,
• one for the KFS filesystem,
• Amazon’s S3, and
• HDFS (the Hadoop Distributed Filesystem)
HBase can persist to any of these implementations.
Hbase v/s RDBMS
• As described previously, HBase is a distributed, column-oriented
data storage system.
• It picks up where Hadoop left off by providing random reads and
writes on top of HDFS.
• It has been designed from the ground up with a focus on scale in
every direction:
– tall in numbers of rows (billions),
– wide in numbers of columns (millions), and
– to be horizontally partitioned and replicated across thousands of
commodity nodes automatically.
• The table schemas mirror the physical storage, creating a system
for efficient data structure serialization, storage, and retrieval.
• The burden is on the application developer to make use of this
storage and retrieval in the right way.
Hbase v/s RDBMS
• Typical RDBMSs are
– fixed-schema,
– row-oriented databases with ACID properties and a
– sophisticated SQL query engine.
• The emphasis is on
– strong consistency,
– referential integrity,
– abstraction from the physical layer, and
– complex queries through the SQL language.
• You can easily
– create secondary indexes,
– perform complex inner and outer joins,
– count, sum, sort, group, and page your data across a number of
tables, rows, and columns.
Guidelines for designing a schema for
HBase
• Joins
– There is no native database join facility in HBase, but wide
tables can make it so that there is no need for database
joins pulling from secondary or tertiary tables.
– A wide row can sometimes be made to hold all data that
pertains to a particular primary key.
• Row keys
– Take time designing your row key.
– A smart compound key can be used clustering data in ways
amenable to how it will be accessed.
– If your keys are integers, use a binary representation
rather than persist the string version of a number—it
consumes less space.
Hbase Use Case
• Webtable: a table of crawled web pages and their attibutes
(such as language and MIME type) keyed by the web page
URL.
• The webtable is large with row counts that run into the
billions.
• Batch analytic and parsing MapReduce jobs are
continuously run against the webtable deriving statistics
and adding new columns of MIME type and parsed text
content for later indexing by a search engine.
• Concurrently, the table is randomly accessed by crawlers
running at various rates updating random rows while
random web pages are served in real time as users click on
a website’s cached-page feature.
Hbase in operation
• HBase, internally, keeps special catalog tables named -ROOT- and .META. within
which it maintains
• the current list,
• state,
• recent history, and
• location of all regions
afloat on the cluster.
• The -ROOT- table holds the list of .META. table regions.
• The .META. Table holds the list of all user-space regions.
• Entries in these tables are keyed using the region’s start row.
• Row keys, as noted previously, are sorted so finding the region that hosts a
particular row is a matter of a lookup to find the first entry whose key is greater
than or equal to that of the requested row key.
• As regions transition—are split, disabled/ enabled, deleted, redeployed by the
region load balancer, or redeployed due to a regionserver crash—the catalog
tables are updated so the state of all regions on the cluster is kept current.
Hbase in operation
• Fresh clients connect to the ZooKeeper cluster
first to learn the location of -ROOT-.
• Clients consult -ROOT- to elicit the location of the
.META. region whose scope covers that of the
requested row.
• The client then does a lookup against the
found .META. Region to figure the hosting user-
space region and its location.
• Thereafter the client interacts directly with the
hosting regionserver.
Hbase in operation
• Clients cache all they learn traversing -ROOT- and
.META. caching locations as well as user-space region
start and stop rows so they can figure hosting regions
themselves without having to go back to the .META.
table.
• Clients continue to use the cached entry as they work
until there is a fault.
• When this happens—the region has moved—the client
consults the .META. again to learn the new location.
• If, in turn, the consulted .META. region has moved,
then -ROOT- is reconsulted.
Hbase in operation: Writes
• Writes arriving at a regionserver are first appended to a commit log
and then are added to an in-memory cache.
• When this cache fills, its content is flushed to the filesystem.
• The commit log is hosted on HDFS, so it remains available through a
regionserver crash.
• When the master notices that a regionserver is no longer
reachable, it splits the dead regionserver’s commit log by region.
• On reassignment, regions that were on the dead regionserver,
before they open for business, pick up their just-split file of not yet
persisted edits and replay them to bring themselves up-to-date
with the state they had just before the failure.
Hbase in operation: Reads
END
• References:
– http://en.wikipedia.org/wiki/Column-oriented_DB
MS
– http://hadoop.apache.org/hbase/
– http://jimbojw.com/wiki/index.php?title=Underst
anding_Hbase_and_BigTable

You might also like