Building My Own Database in C++ — Part 2: Querying the Database
Table of Contents
Defining a schema
So far, our storage engine has treated tuples as raw bytes. While this is sufficient for storing records on disk, a database also needs to understand what those bytes represent. For example, a table might contain an id column containing integers and a name column containing strings.
To represent this information, we introduce three abstractions: Value, Column, and Schema. A Value represents an individual database value, a Column describes one field in a table, and a Schema describes the structure of an entire table.
We first define the types supported by our database using an enum class. For now, the database supports integers, booleans, larger integers, and variable-length strings.
1enum class TypeId
2 {
3 INTEGER,
4 BIGINT,
5 BOOLEAN,
6 VARCHAR
7 }; The VARCHAR type represents a variable-length string. Unlike an INTEGER, which always occupies a fixed number of bytes, a VARCHAR value can contain a different number of characters depending on the value. For example, "Alice" and "Christopher" have different lengths.
A Value represents one actual value stored in a tuple. The class keeps track of the value's type and stores the corresponding data. Factory methods are used to make constructing typed values explicit.
1namespace db
2 {
3 class Value
4 {
5 public:
6 static Value Integer(int32_t value) // factory methods to create a value
7 {
8 Value v;
9 v.type_ = TypeId::INTEGER;
10 v.int_value_ = value;
11 return v;
12 }
13
14 static Value BigInt(int64_t value)
15 {
16 Value v;
17 v.type_ = TypeId::BIGINT;
18 v.bigint_value_ = value;
19 return v;
20 }
21
22 static Value Boolean(bool value)
23 {
24 Value v;
25 v.type_ = TypeId::BOOLEAN;
26 v.bool_value_ = value;
27 return v;
28 }
29
30 static Value Varchar(const std::string &value)
31 {
32 Value v;
33 v.type_ = TypeId::VARCHAR;
34 v.string_value_ = value;
35 return v;
36 }
37
38 TypeId type() const
39 {
40 return type_;
41 }
42
43 int32_t as_int() const
44 {
45 return int_value_;
46 }
47
48 int64_t as_bigint() const
49 {
50 return bigint_value_;
51 }
52
53 bool as_bool() const
54 {
55 return bool_value_;
56 }
57
58 const std::string &as_string() const
59 {
60 return string_value_;
61 }
62
63 private:
64 TypeId type_;
65
66 int32_t int_value_ = 0;
67 int64_t bigint_value_ = 0;
68 bool bool_value_ = false;
69 std::string string_value_;
70 };
71
72 } For example, we can now create actual database values instead of manually constructing their serialized byte representation. The Integer() and Varchar() methods are static factory methods. Because they are static, they can be called directly using the class name without first creating a Value object. Each method constructs a Value, assigns its type, stores the corresponding data, and returns it.
Next, we need a way to describe what each field in a table represents. A Column stores the column's name and its type.
1namespace db
2 {
3 class Column
4 {
5 public:
6 Column(
7 std::string name,
8 TypeId type)
9 : name_(std::move(name)),
10 type_(type)
11 {
12 }
13
14 const std::string &name() const
15 {
16 return name_;
17 }
18
19 TypeId type() const
20 {
21 return type_;
22 }
23
24 private:
25 std::string name_;
26 TypeId type_;
27 };
28 }A column can now be defined using its name and type.
Finally, a Schema combines multiple columns to describe the structure of an entire table. The schema stores the columns in order, allowing each column to be identified by its index.
1namespace db
2 {
3
4 class Schema
5 {
6 public:
7 explicit Schema(std::vector<Column> columns)
8 : columns_(std::move(columns))
9 {
10 }
11
12 size_t column_count() const
13 {
14 return columns_.size();
15 }
16
17 const Column &column(size_t index) const
18 {
19 return columns_[index];
20 }
21
22 int column_index(const std::string &name) const
23 {
24 for (int i = 0; i < columns_.size(); i++)
25 {
26 if (columns_[i].name() == name)
27 {
28 return i;
29 }
30 }
31
32 return -1;
33 }
34
35 private:
36 std::vector<Column> columns_;
37 };
38 } At this point, our database can represent both the structure of a table and the values stored within its rows. The Schema describes what columns exist and their types, while Value represents the actual data stored in those columns. The next step is to connect these logical representations to our Tuple class so that values can be serialized into the raw bytes stored by TablePage.