You are on page 1of 80

Syntax-Directed Translation

1
Syntax-Directed Translation
 We associate information with the programming language constructs by
attaching attributes to grammar symbols.
 Values of these attributes are evaluated by the semantic rules associated with the
production rules.
 Evaluation of these semantic rules:
• may generate intermediate codes
• may put information into the symbol table
• may perform type checking
• may issue error messages
• may perform some other activities

 An attribute may hold almost any thing.


• a string, a number, a memory location, a complex record.

2
Syntax-Directed Definitions & Translation Schemes

 When we associate semantic rules with productions, we use two


notations:

A. Syntax-Directed Definitions:
– High level specifications

– Hide implementation details such as order of evaluation of semantic actions.

B. Translation Schemes:
– Low level specification

– Give a little bit information about implementation details such as order of evaluation of
semantic actions associated with a production rule

3
Syntax-Directed Translation
• Conceptually with both the syntax directed translation and translation
scheme we can
– Parse the input token stream
– Build the parse tree
– Traverse the tree to evaluate the semantic rules at the parse tree nodes.

Input string parse tree dependency graph evaluation order for


semantic rules

Conceptual view of syntax directed translation

4
Syntax-Directed Definitions
 A syntax-directed definition is a generalization of a context-free
grammar in which:
– Each grammar symbol is associated with a set of attributes.

– Set of attributes for a symbol is partitioned into synthesized and inherited.

– Each production rule is associated with a set of semantic rules.

 The value of an attribute at a parse tree node is defined by the semantic rule
associated with a production at that node.
 The value of a synthesized attribute at a node is computed from the values of
attributes at the children in that node of the parse tree
 The value of an inherited attribute at a node is computed from the values of
attributes at the siblings and parent of that node of the parse tree 5
Syntax-Directed Definitions
Examples:
Synthesized attribute : E→E1+E2 { E.val =E1.val + E2.val}
Inherited attribute :A→XYZ {Y.val = 2 * A.val}

 Semantic rules set up dependencies between attributes which can be


represented by a dependency graph.
 Dependency graph determines the evaluation order of semantic rules.

 Evaluation of a semantic rule defines the value of an attribute.

 Semantic rule may also have some side effects such as printing a
value.
6
Annotated Parse Tree

 A parse tree showing the values of attributes at each node is called


an annotated parse tree.
 Values of Attributes in nodes of annotated parse-tree are either,
• initialized to constant values or by the lexical analyzer.

• determined by the semantic-rules.

 The process of computing the attributes values at the nodes is called


annotating (or decorating) of the parse tree.
 The order of these computations depends on the dependency graph
induced by the semantic rules.
7
Syntax-Directed Definition
 In a syntax-directed definition, each production A→α is associated
with a set of semantic rules of the form:

b=f(c1,c2,…,cn)

where f is a function and b can be one of the followings:


– b is a synthesized attribute of A and c1,c2,…,cn are attributes of the
grammar symbols in the production ( A→α ).
OR
– b is an inherited attribute one of the grammar symbols in α (on the right
side of the production), and c1,c2,…,cn are attributes of the grammar
symbols in the production ( A→α ). 8
Attribute Grammar

 A semantic rule b=f(c1,c2,…,cn) indicates that the attribute b depends on

attributes c1,c2,…,cn.

 In a syntax-directed definition, a semantic rule may just evaluate


a value of an attribute or it may have some side effects such as
printing values.
 An attribute grammar is a syntax-directed definition in which the
functions in the semantic rules cannot have side effects (they can only
evaluate values of attributes).

9
Syntax-Directed Definition -- Example

Production Semantic Rules


L→En print(E.val)
E → E1 + T E.val = E1.val + T.val
E→T E.val = T.val
T → T1 * F T.val = T1.val * F.val
T→F T.val = F.val
F→(E) F.val = E.val
F → digit F.val = digit.lexval
 Symbols E, T, and F are associated with a synthesized attribute val.
 The token digit has a synthesized attribute lexval (it is assumed that it is evaluated by
the lexical analyzer).
 Terminals are assumed to have synthesized attributes only. Values for attributes of
terminals are usually supplied by the lexical analyzer.
 The start symbol does not have any inherited attribute unless otherwise stated.
10
Synthesized attributes

 Synthesized attributes depend only on the attributes of children. They are the
most common attribute type.

 If a SDD has synthesized attributes ONLY, it is called a S-ATTRIBUTED


DEFINITION.

11
Inherited attributes

 An inherited value at a node in a parse tree is defined in terms of attributes at the


parent and/or siblings of the node.
 Convenient way for expressing the dependency of a programming language construct
on the context in which it appears.
 We can use inherited attributes to keep track of whether an identifier appears on the
left or right side of an assignment to decide whether the address or value of the
assignment is needed.
 Example: The inherited attribute distributes type information to the various identifiers
in a declaration.

12
Inherited attributes

 An inherited attribute is defined in terms of the attributes of the node’s


parents and/or siblings.

 Inherited attributes are often used in compilers for passing contextual


information forward, for example, the type keyword in a variable
declaration statement.

13
Syntax-Directed Definition – Inherited Attributes

Production Semantic Rules


D→TL L.in = T.type
T → int T.type = integer
T → real T.type = real
L → L1,id L1.in = L.in, addtype(id.entry,L.in)
L → id addtype(id.entry,L.in)

1. Symbol T is associated with a synthesized attribute type.

1. Symbol L is associated with an inherited attribute in.

14
Annotated parse tree
Input: real p,q,r annotated parse
tree
parse tree D
D

T L T.type=real L1.in=real

real L , id3 real L1.in=real , id3

L , id2 L1.in=real , id2

15
Evaluating an SDD at the Nodes of a Parse Tree

 First construct the parse tree

 Using the rules to evaluate all of the attributes at each of the nodes of
the parse tree
 A parse tree showing the value(s) of the attribute(s) is called an
annotated parse tree

16
This SDD uses a combination of synthesized and inherited attributes.
A.s (head) is defined in terms of B.i (body nonterminal). Hence, it is
synthesized.
B.i (body non-terminal) is defined in terms of A.s (head). Hence, it is
inherited.
There exists a circular dependency between their evaluations.
In practice, subclasses of SDDs required for our purpose do have an order.

17
Syntax-Directed Definition -- Example

Production Semantic Rules


L→En print(E.val)
E → E1 + T E.val = E1.val + T.val
E→T E.val = T.val
T → T1 * F T.val = T1.val * F.val
T→F T.val = F.val
F→(E) F.val = E.val
F → digit F.val = digit.lexval

18
Syntax-Directed Definition – Inherited Attributes

Production Semantic Rules


D→TL L.in = T.type
T → int T.type = integer
T → real T.type = real
L → L1, id L1.in = L.in, addtype(id.entry,L.in)
L → id addtype(id.entry,L.in)

19
Annotated parse tree
Input: real p,q,r annotated parse
tree
parse tree D
D

T L T.type=real L1.in=real

real L , id3 real L1.in=real , id3

L , id2 L1.in=real , id2

20
Annotated Parse Tree for 3*5

21
Evaluation Order for SDD’s

 Dependency graphs are useful tools for determining an evaluation order

for the attribute instances in a given parse tree

 Annotated parse tree – values of attributes

 Dependency graphs – determine how those values can be computed

22
 A dependency graph shows the flow of information among the
attribute instances in a particular parse tree; an edge from one
attribute instance to another means that the value of the first is needed
to compute the second.
 Edges express constraints implied by the semantic rules.
 Each attribute is associated to a node

 If a semantic rule associated with a production p defines the value of synthesized


attribute A.b in terms of the value of X.c, then graph has an edge from X.c to A.b
 If a semantic rule associated with a production p defines the value of inherited
attribute B.c in terms of value of X.a, then graph has an edge from X.a to B.c

23
Dependency Graph Construction

for each node n in the parse tree do

for each attribute a of the grammar symbol at n do

construct a node in the dependency graph for a

for each node n in the parse tree do

for each semantic rule b = f(c1,…,cn) associated with production used at n do

for i= 1 to n do

construct an edge from the node for ci to the node for b

24
Dependency Graph Construction
• Example
• Production Semantic Rule
E→E1 + E2 E.val = E1.val + E2.val

E . val

E1. val + E2 . Val


• E.val is synthesized from E1.val and E2.val
• The dotted lines represent the parse tree that is not part of the
dependency graph.

25
Dependency Graph – 3*5

26
Dependency Graph
D→TL L.in = T.type
T → int T.type = integer
T → real T.type = real
L → L1 id L1.in = L.in,
addtype(id.entry,L.in)

L → id addtype(id.entry,L.in)

27
Ordering the Evaluation of Attributes
 A dependency graph characterizes the possible order in which we can evaluate
the attributes at various nodes of a parse tree.
 If there is an edge from node M to N, then attribute corresponding to M first
be evaluated before evaluating N.

 Thus the only allowable orders of evaluation are N1, N2, ..., Nk such that if

there is an edge from Ni to Nj then i < j. Such an ordering embeds a directed


graph into a linear order, and is called a topological sort of the graph.
 If there is any cycle in the graph, then there are no topological sorts; that is,
there is no way to evaluate the SDD on this parse tree.
 If there are no cycles, however, then there is always at least one topological
28

sort.
Evaluation Order

• a4=real;
• a5=a4;
• addtype(id3.entry,a5);
• a7=a5;
• addtype(id2.entry,a7);
• a9=a7;
• addtype(id1.entry,a5);

29
Evaluating Semantic Rules
 Parse Tree methods
– At compile time evaluation order obtained from the topological sort of
dependency graph.
– Fails if dependency graph has a cycle
 Rule Based Methods
– Semantic rules analyzed by hand or specialized tools at compiler construction
time
– Order of evaluation of attributes associated with a production is pre-determined at
compiler construction time
 Oblivious Methods
– Evaluation order is chosen without considering the semantic rules.
– Restricts the class of syntax directed definitions that can be implemented.
– If translation takes place during parsing order of evaluation is forced by parsing
method.

30
S-attributed definition
 A syntax directed translation that uses synthesized attributes exclusively is said to be
a S-attributed definition.
 A parse tree for a S-attributed definition can be annotated by evaluating the semantic
rules for the attributes at each node, bottom up from leaves to the root.
 Evaluation is simple using post-order traversal.

postorder(N)
{
for (each child C of N, from the left) postorder(C);
evaluate attributes associated with node N;
}

31
L-Attributed Definitions
 Dependency-graph edges can go from left to right, but not from right to
left(hence "L-attributed").
 Each attribute must be either Synthesized, or Inherited, but with the rules limited
as follows.
• Suppose that there is a production A X1 X2 • • • Xn, and that there is an
inherited attribute Xi.a computed by a rule associated with this production.
Then the rule may use only:
(a) Inherited attributes associated with the head A.

(b) Either inherited or synthesized attributes associated with the occurrences of symbols X1

X2 • • • Xi-1 located to the left of Xi

(c) Inherited or synthesized attributes associated with this occurrence of Xi itself, but only
in such a way that there are no cycles in a dependency graph formed by the attributes
of this Xi.
Example for L-attributed SDD

33
A Definition which is not L-Attributed
Depth-first traversal
L-attributed definitions are the set of SDDs whose attributes can
be evaluated in a DEPTH-FIRST traversal of the parse tree.

dfvisit( node n ) {
for each child m of n, in left-to-right order, do {
evaluate the inherited attributes of m
dfvisit( m )
}
evaluate the synthesized attributes of n
}

35
Semantic Rules with Controlled Side Effects

 Side Effects: Evaluation of semantic rules may generate intermediate codes,

may put information into the symbol table, may perform type checking and

may issue error messages. These are known as side effects.

 In practice translation involves side effects.

– a desk calculator might print a result;

– a code generator might enter the type of an identifier into a symbol table.

• Attribute grammars have no side effects and allow any evaluation order

consistent with dependency graph. 36


Ways to Control Side Effects

 Permit incidental side effects that do not constrain attribute evaluation.


– Permit side effects when attribute evaluation based on any topological sort
of the dependency graph produces a "correct" translation, where "correct"
depends on the application.

 Constrain the allowable evaluation orders, so that the same translation


is produced for any allowable order.
– The constraints can be thought of as implicit edges added to the
dependency graph.

37
Example - incidental side effect
Modify the desk calculator to print a result. Instead of the rule L.val
= E.val, which saves the result in the synthesized attribute L.val

Semantic rules that are executed for their side effects, such as
print(E.val), will be treated as the definitions of dummy synthesized
attributes associated with the head of the production.

The modified SDD produces the same translation under any topological
sort, since the print statement is executed at the end, after the result is
computed into E.val.

38
 Productions 4 and 5 have a rule in which a function addType is called
with two arguments:
1. id.entry, a lexical value that points to a symbol-table object, and
2. L.inh, the type being assigned to every identifier on the list.
 The function addType properly installs the type L.inh as the type of the
represented identifier.
 This side effect, adding the type info to the table, does not affect the
evaluation order.

39
 Numbers 1 through 10 represent the nodes of the dependency graph.
 Nodes 1, 2, and 3 represent the attribute entry associated with each of
the leaves labeled id.
 Nodes 6, 8, and 10 are the dummy attributes that represent the
application of the function addType to a type and one of these entry
values.

40
Applications of Syntax-Directed Translation

 Type checking

 Intermediate Code Generation

 Construction of Syntax Trees

41
Syntax Trees
Syntax-Tree
– an intermediate representation of the compiler’s input.

– A condensed form of the parse tree.

– Syntax tree shows the syntactic structure of the program while


omitting irrelevant details.
– Operators and keywords are associated with the interior nodes.

Syntax directed translation can be based on syntax tree as well as parse


tree.

42
Parse Tree vs Syntax Tree

43
Syntax Tree-Examples
Expression: Statement:
+ if B then S1 else S2
if - then - else

5 *
B S1 S2
3 4
• Leaves: identifiers or constants • Node’s label indicates what kind
• Internal nodes: labelled with of a statement it is
operations • Children of a node correspond to
• Children: of a node are its the components of the statement
operands
44
Construction of Syntax Trees
 Nodes of a syntax tree can be implemented by objects with a suitable
number of fields.
 Each object will have an op field that is the label of the node.
 The objects will have additional fields as follows:

– If the node is a , an additional field holds the lexical value for the . A
constructor function Leaf(op, val) creates a object. Alternatively, if nodes
are viewed as records, then returns a pointer to a new record for a .
– If the node is an interior node, there are as many additional fields as the
node has children in the syntax tree. A constructor function Node takes
two or more arguments: Node(op,c1,c2,... ,ck) creates an object with first
field op and k additional fields for the k children
45
46
Constructing Syntax Tree for Expressions
Example: a-4+c

1. p1:= new Leaf (id, entry-a);


2. p2:= new Leaf (num,4);
3. p3 := new Node(‘-’,p1,p2) id
4. p4:= new Leaf (id, entry-c);
5. p5:= new Node(‘+’,p3,p4); to entry for c
id num 4
• The tree is constructed bottom
up. to entry for a

47
48
int [2][3]

49
Syntax-Directed Translation Schemes

 Translation schemes are another way to describe syntax-directed translation.

 A translation scheme is a context-free grammar in which:

– attributes are associated with the grammar symbols and

– semantic actions enclosed between braces {} are inserted within the right sides of
productions.

 If braces are needed as grammar symbols, then we quote them.

 Translation schemes describe the order and timing of attribute computation.

– specify when, during the parse, attributes should be computed


Syntax-Directed Translation Scheme - Example
A Translation Scheme Example
The depth first traversal of the parse tree (executing the semantic actions in that order)
will produce the postfix representation of the infix expression.
Postfix Translation Schemes
 The simplest SDD implementation occurs when we can parse the grammar bottom-up
and the SDD is S-attributed.

 We can construct an SDT in which each action is placed at the end of the
production and is executed along with the reduction of the body to the head of that
production.

 SDT's with all actions at the right ends of the production bodies are called postfix
SDT's.
Parser-Stack Implementation of Postfix SDT's

 Postfix SDT's can be implemented during LR parsing by executing the

actions when reductions occur.

 The attribute(s) of each grammar symbol can be put on the stack in a place

where they can be found during the reduction.

 The best plan is to place the attributes along with the grammar symbols in

records on the stack itself.


 The parser stack contains records with a field for a grammar symbol (or parser state)
and, below it, a field for an attribute.

 The three grammar symbols X YZ are on top of the stack; perhaps they are about to be
reduced according to a production like A XYZ.

 X.x as the one attribute of X, and so on.

 If one or more attributes are of unbounded size - say, they are character strings - then
it would be better to put a pointer to the attribute's value in the stack record and
store the actual value in some larger, shared storage area that is not part of the stack.
Translation Schemes for S-attributed Definitions
 If the attributes are all synthesized, and the actions occur at the ends of the
productions, then we can compute the attributes for the head when we reduce
the body to the head.
 If we reduce by a production such as A XYZ, then we have all the attributes
of X, Y, and Z available, at known positions on the stack.
 After the action, A and its attributes are at the top of the stack, in the position
of the record for X.

E → E1 + T E.val = E1.val + T.val



E → E1 + T { E.val = E1.val + T.val }
 Semantic Actions during Parsing
– when shifting
• push the value of the terminal on the semantic stack
– when reducing
• pop k values from the semantic stack, where k is the number of symbols on
production’s RHS
• push the production’s value on the semantic stack
Bottom-Up Evaluation -- Example
Input state val semantic rule
5+3*4n - -
+3*4n 5 5
+3*4n F 5 F → digit
+3*4n T 5 T→F
+3*4n E 5 E→T
3*4n E+ 5-
*4n E+3 5-3
*4n E+F 5-3 F → digit
*4n E+T 5-3 T→F
4n E+T* 5-3-
n E+T*4 5-3-4
n E+T*F 5-3-4 F → digit
n E+T 5-12 T → T1 * F
n E 17 E → E1 + T
En 17- L→En
L 17
SDT's With Actions Inside Productions

 An action may be placed at any position within the body of a production.

 It is performed immediately after all symbols to its left are processed.

 If we have a production B  X {a} Y, the action a is done after we have


recognized X (if X is a terminal) or all the terminals derived from X (if X is a
nonterminal).

 If the parse is bottom-up, then we perform action a as soon as this occurrence


of X appears on the top of the parsing stack.

 If the parse is top-down, we perform a just before we attempt to expand this


occurrence of Y (if Y a nonterminal) or check for Y on the input (if Y is a
terminal).
SDT for Infix to Prefix
Eliminating Left Recursion From SDT's

 Grammar with left recursion cannot be parsed using top-down parser.

 In case of SDT we treat action as terminal symbol in the production.

 The "trick" for eliminating left recursion is to take two productions replace
them by productions that generate the same strings using a new nonterminal R
(for "remainder") of the first production:
Eliminating Left Recursion - Example
Eliminating Left Recursion - Example
SDT's for L-Attributed Definitions

 It is necessary that the underlying grammar can be parsed top-down

rules to modify L-attributed SDD to SDT

– Embed the action for computing inherited attributes for a nonterminal A

immediately before A in production body. If several inherited attributes for A

depend on one another in an acyclic fashion, order the evaluation of attributes so

that those needed first are computed first.

– Place the action for computing synthesized attribute for the head at the end of the

body of production
Example 1
 This example is motivated by languages for typesetting mathematical
formulas.
 Eqn is an early example of such a language.
 It illustrates how the techniques of compiling can be used in language
processing for applications other than programming languages.
 Corresponding to the following four productions, a box can be either
1. Two boxes, connected, with the first, B1, to the left of the other, B2.
2. A box and a subscript box. The second box appears in a smaller size, lower, and
to the right of the first box.
3. A parenthesized box, for grouping of boxes and subscripts.
4. A text string, that is, any string of characters.
The function new generates new labels.
The variables L1 and L2 hold labels that we need in the code.
We use || as the symbol for concatenation of intermediate-code fragments.The value of S.code thus begins with the label L1, then the code for condition C, another label L2, and the code for S1.
S1

Example 2

• The function new generates new labels.

• The variables L1 and L2 hold labels that we need in the code.

• We use || as the symbol for concatenation of intermediate-code fragments.


We use the following attributes to generate the proper intermediate code:
1. The inherited attribute S.next labels the beginning of the code that must be executed
after S is finished.

2. The synthesized attribute S.code is the sequence of intermediate-code steps that


implements a statement S and ends with a jump to S.next.

3. The inherited attribute C.true labels the beginning of the code that must be executed
if C is true.

4. The inherited attribute C.false labels the beginning of the code that must be executed
if C is false.

5. The synthesized attribute C.code is the sequence of intermediate-code steps that


implements the condition C and jumps either to C.true or to C.false, depending on
whether C is true or false.
Implementing L-attributed SDD’s
1. Build the parse tree and annotate - noncircular SDD.

2. Build the parse tree, add actions, and execute the actions in preorder - L-attributed
definition.

3. Use a recursive-descent parser with one function for each nonterminal - function
for nonterminal A receives the inherited attributes of A as arguments and returns the
synthesized attributes of A.

4. Generate code on the fly, using a recursive-descent parser

5. Implement an SDT in conjunction with an LL-parser - attributes are kept on the


parsing stack, and the rules fetch the needed attributes from known locations on the
stack.

6. Implement an SDT in conjunction with an LR-parser


Translation During Recursive-Descent Parsing
 We can extend the Recursive descent parser into a translator as follows:

a) The arguments of function A are the inherited attributes of nonterminal A.

b) The return-value of function A is the collection of synthesized attributes of nonterminal


A.

 In the body of function A, we need to both parse and handle attributes:

a) Decide upon the production used to expand A.

b) Check that each terminal appears on the input when it is required. We shall assume that
no backtracking is needed, but the extension to recursive- descent parsing with
backtracking can be done by restoring the input position upon failure

c) Preserve, in local variables, the values of all attributes needed to compute inherited
attributes for nonterminals in the body or synthesized attributes for the head nonterminal.

d) Call functions corresponding to nonterminals in the body of the selected production,


providing them with the proper arguments.
Implementing while-statements with recursive-
descent parser

Swhile ( { L1=new();L2=new();C.false=S.next;C.true=L2; }
C ) { S1.next==L1; }
S1 { S.code=label||L1||C.code||label||L2|| S1.code; }
Recursive-descent typesetting of boxes
On-The-Fly Code Generation
 The construction of long strings of code that are attribute values – undesirable - time it
could take to copy or move
 we can instead incrementally generate pieces of the code into an array or output file
by executing actions in an SDT.

 The elements we need to make this technique work are:

1. There is a main attribute for one or more nonterminals, - e.g.: S.code and C.code

2. The main attributes are synthesized.

3. The rules that evaluate the main attribute(s) ensure that

a) The main attribute is the concatenation of main attributes of nonterminals


appearing in the body of the production involved, perhaps with other elements that
are not main attributes, - e.g.: labels L1 and L2.

b) The main attributes of nonterminals appear in the rule in the same order as the
nonterminals themselves appear in the production body.
On-the-fly recursive-descent code generation for
while-statements
L-Attributed SDD's and LL Parsing
 We can then perform the translation during LL parsing by extending the parser stack
to hold actions and certain data items needed for attribute evaluation.

 Parser stack will hold


– records representing terminals and nonterminals

– action-records representing actions to be executed and

– Synthesize-records to hold the synthesized attributes.

 We use the following two principles to manage attributes on the stack:


– The inherited attributes of a nonterminal A are placed in the stack record. The code to
evaluate these attributes will usually be represented by an action-record immediately
above the stack record for A

– The synthesized attributes for a nonterminal A are placed in a separate synthesize-record


that is immediately below the record for A on the stack.
While-statement Translation sequences

After the action above C is performed

Expansion of S according to the while-statement production


While-statement Translation sequences

Expansion of S with synthesized attribute constructed on the stack


Bottom-Up Parsing of L-Attributed SDD’s(LR)

 This has 3 parts


1. Start with the SDT, which places embedded actions before each nonterminal to
compute its inherited attributes and an action at the end of the production to
compute synthesized attributes.
2. Introduce into the grammar a marker nonterminal in place of each embedded
action. Each such place gets a distinct marker, and there is one production for
any marker M, namely M Є.

3. Modify the action a if marker nonterminal M replaces it in some production


A {a} , and associate with M Є an action a' that
a) Copies, as inherited attributes of M, any attributes of A or symbols of  that action a
needs.

b) Computes attributes in the same way as a, but makes those attributes be synthesized
attributes of M.
Example

ABC
A{ B.i = f(A.i);} BC
LR parsing stack after reduction of Є to M

Stack just before reduction of the while-production body to S

You might also like