You are on page 1of 3

2017 IEEE 17th International Conference on Advanced Learning Technologies

Automated Personalized Assessment of


Computational Thinking MOOC Assignments

Hasan M. Jamil
Department of Computer Science
University of Idaho, USA
jamil@uidaho.edu

Abstract—One of the major hurdles toward automatic se- The encouraging fact is that there have been significant
mantic understanding of computer programs is the lack of progress in several areas of computer science that we believe
knowledge about what constitutes functional equivalence of code can be leveraged and assembled together to build effective
segments. We postulate that a sound knowledgebase can be and smart cyber systems for online teaching, tutoring and
used to deductively understand code segments in a hierarchical assessment of entry level computing classes, and other STEM
fashion by first de-constructing a code and then reconstructing it subjects. In our vision, such a system will complement a
from elementary knowledge and equivalence rules of elementary
code segments. The approach can also be engineered to produce
human instructor or mentor, and take on the role of a human
computable programs from conceptual and abstract algorithms observer to monitor students in real time and detect where she
as an inverse function. In this paper, we introduce the core idea is making a mistake in her coding exercise, and immediately
behind the MindReader online assessment system that is able offer assistance by providing diagnostic comments and helpful
to understand a wide variety of elementary algorithms students pointers that most likely will cure the error [5].
learn in their entry level programming classes such as Java, C++
and Python. The MindReader system is able to assess student In this paper, we introduce a novel prototype online sys-
assignments and guide them how to develop correct and better tem for tutoring and assessment, called the MindReader, for
code in real time without human assistance. high school and freshman college students to aid learning
programming languages. We develop necessary computational
Keywords—Authentic assessment; computational thinking; au-
tomated assessment; computer programming; program equivalence;
technologies to advance the science of computer program
understanding needed for digital tutoring, and online real time
assessment of programming assignments. This system is aimed
at complementing human instructors at a more massive scale
I. I NTRODUCTION fully automatically. For the want of space and brevity, however,
we highlight only the salient features of MindReader and refer
Technological advances and economic realities are prompt-
readers to [6] for a more detailed discussion.
ing a shift in the way we learn, teach and deliver instructions
to train our labor force. Tech savvy younger generation today
find personalized online systems engaging and useful and are II. R ELATED R ESEARCH
welcoming online and digital learning in all three settings –
formal or institutional, blended, and self-paced and non-formal While some progress has been made in online instruction
learning spaces. The expectation is that online systems will delivery, as well as in creating exciting learning environments,
overcome much of the hurdles we face in formal education real time assessment and tutoring of STEM subjects online
systems and will complement it in a larger way. Although some still remain at its infancy. Most often than not, these two areas
skepticism exists [1], the excitement around Massive Open rely largely on human interaction or MCQ tests, effectiveness
Online Courses (MOOC) and more institutional approach to of which are still being debated in general [7], and in STEM
digital learning using BBLearn1 that are at the two ends of a settings [8] and for computer programming classes [9] in
spectrum, are fueled by these promises. All online universities, particular. The general consensus appear to be that MCQ tests
academies and institutes such as Coursera2 , KhanAcademy3 , are great tools for formative and diagnostic assessments but
and MITOpenCourseware4 are then immediately faced with for summative assessment, tests such as authentic assessment
problems in three axes – content delivery, teaching, and are more appropriate, particularly in computing courses.
tutoring and assessment, much like the traditional systems do.
The challenges in designing smart cyber systems for tutor-
They also grapple with enrollment and coverage, retention,
ing and summative assessment are manifold. Ideally, a tutoring
cost, teaching effectiveness, and so on much like their formal
or assessment system should not rely on a specific procedure
counterparts. New learning environments such as immersive,
for establishing correctness of a proof in mathematics, for
game-based, blended, personalized, self-paced, social, peer,
example. Rather the logical argument in any order must be the
and pair learning have been proposed to combat them, the
basis. Such an assumption rules out most of the current ap-
effectiveness studies of which are ongoing [2], [3], [4].
proaches to establishing correspondence of a student response
1 http://www.blackboard.com/ to a known solution. A few automated systems have attempted
2 https://www.coursera.org/ to capture this spirit in subjects such as mathematics [10]
3 https://www.khanacademy.org/computing/computer-programming and physics [11] education with extremely narrow success. In
4 https://ocw.mit.edu/courses/intro-programming/ computer science, the success has been mixed [12], [13].

2161-377X/17 $31.00 © 2017 IEEE 261


DOI 10.1109/ICALT.2017.147
To assess programming assignments, usually understanding semantically as opposed to syntactic matching using PDGs.
the code semantically is required, and for a machine it would In CDGs, concepts are defined hierarchically, and terminal
essentially mean determining the functional equivalence of a nodes are either declaration, or base computable concepts such
reference solution and the student solution, which is theoreti- as print, decide, for loop, while loop, and so on such that
cally hard – deciding functional equivalence of two programs all variables needed for the base concepts are also in the
in general is NP-complete [14], and only in limited instances CDG. CDGs can be simple or complex. A simple CDG is
and for special classes of programs we are able to do so a connected and directed acyclic graph of base concepts and
[15]. Undeterred by this weakness, researchers took a different declarations, while a complex CDG is a forest of simple CDGs
route and tried to assess correctness of programs by various and CDGs involving connected concept nodes. A CDG graph
means so that the method can be used in learning exercises represents the precedence relationship among the concepts.
and online settings [16] but faced complexity barriers of a The claim is that the matching of two CDGs in this manner
different nature [17]. Other approaches used test data to assess have a much higher likelihood of determining functional and
correctness [18] to match with known outcomes and “assume” semantic equivalence of two code segments necessary for
correctness. We, however, are not aware of a system capable grading assignments, and offering tutoring help. We illustrate
of tutoring or assessing computer programs fully automatically the idea using a simple problem of averaging a list of values
and comprehensively. in C++.
Technically, a CDG is a graph ⟨𝑁, ≺⟩ of a set of nodes
III. CDG S : H IERARCHICAL C ONCEPT S TRUCTURE 𝑁 = ∪𝑖 𝑛𝑖 ⊆ 𝑉 and a precedence relation ≺= ∪𝑗 𝑒𝑗 ⊆ 𝐸,
The program dependence graph (PDG) [19] based match- where 𝑉 is the set of all possible concepts, and 𝐸 is all
ing approach to determine code equivalence for the purpose of pairs 2𝑉 ×𝑉 . In CDGs, a concept such as a counter loop
grading programming assignments is too simplistic although is replaceable with another set of interconnected concepts
such approaches have been narrowly effective in detecting that performs identical function – a for loop or while loop
code clones [20] and plagiarized codes [21]. In particular, implementation of a counter loop. In other words, we can
such techniques call for a complete enumeration of all possible assume that a counter loop can be implemented in two possible
solutions for every assignment, a largely daunting task, if not ways, and thus establish an equivalence relation. Note that the
impossible. For example, consider an assignment that involves concepts are also hierarchically defined. The concept average
writing a code segment to swap the values of two variables. As is defined as an aggregation of a list of values, followed by
shown in figure 1, a student cannot be penalized if she offered a division by the size of the list. An aggregate on the other
the code segment as a possible solution even though a PDG hand is defined as the summation of the elements of the list
based matching approach will most likely fail to accept it as a inside a counter loop (note node 𝑣4 is part of node 𝑣3 , the
possible solution. If a student offers a more sophisticated but counter loop). Finally a counter loop is defined as a for loop
unanticipated solution instead as shown below, she should be or a while loop. For the student program 𝑃𝑠 below, we can
assigned higher credits, not less, though a PDG based grading transform it to construct a corresponding CDG, and match it
will certainly be ineffective. with the conceptual solution in the knowledgebase even if the
student solution is implemented using a for loop.
void swap(int *i, int *j){
int t = *i; 1: #include <iostream>
*i = *j; 2: void main() {
*j = t; } 3: int k=0, total, size=9, mean, elements[10];
4: while (k<=size) {
5: total=total + elements[k];
6: k++; }
#include <iostream> #include <iostream> 7: std::cout << total/(size+1); }
using std::cout; using std::cout;

int main(){ void swap(int& i, int& j){


int a=27, b=43, t; int t = i; IV. A SSESSMENT AND T UTORING USING M IND R EADER
i = j;
cout << "Before " << a j = t; } The high level architecture in figure 2 depicts Min-
<< " " << b << endl; dReader’s two broad subsystems for two distinct but com-
t = a; int main(){
a = b; int a=27, b=43, t;
plementary functions – tutoring and grading. In MindReader,
b = t; all learners have a profile which includes background, past
cout << "After " << a << cout << "Before " << a lessons, tests and tutoring activities, known problem areas,
" " << b << endl; << " " << b << endl;
return 0; } swap(a, b); and their peer groups. MindReader generates tutorials based
cout << "After " << a << on students’ profile and level of programming competence
" " << b << endl; expected along the lines of the systems such as [22] keeping
return 0; }
in mind that for computational thinking classes, the challenges
Reference Solution Student Solution primarily involve the difficulties in learning the syntax and
understanding the semantics and use of constructs such as
Fig. 1: Equivalent swapping code segments. loops, conditional statements, and simple algorithms [23].
For the purpose of both grading and tutoring, MindReader
assembles the statement structures written by the student into
In this paper, we propose a novel and a more effective possible CDGs using the concept structures in the Concept
approach to matching solutions based on the idea of concept Database according to the rules in Concept Construction rule
dependance graphs (CDGs) in which nodes are matched base with the aim of matching the CDG with one of the known

262
User interface
Test data generator
MindReader Grader Learner profiles Learning objectives Algorithm template Dataflow pattern
graphs database
CDG

Success
MindReader Tutor Dataflow pattern
CDG Visual Editor Concept generator Graph matcher Random tester
matcher

Interactive error
handler Diagnostic generator Concept construction Concept database Crowd/Expert diagnosis
rules

Fig. 2: MindReader architecture: Brown lines for tutoring, purple lines for assessment, and black lines for both system.

templates in the Algorithm Templates. Failure to match CDGs [7] I. Simonova, “Multiple-choice testing: Knowledge, or random choice?”
of the student code and the reference template results into in IEEE EDUCON, Istanbul, Turkey, April 3-5, 2014, pp. 819–823.
a dataflow pattern match using known and random test data [8] J. M. Azevedo, “e-assessment in mathematics courses with multiple-
of the compiled codes. Failure to match flow patterns forces choice questions tests,” in CSEDU , Volume 2, Lisbon, Portugal, 23-25
May, 2015, pp. 260–266.
a diagnostic feedback, but a success indicates a new way of
[9] S. M. Shuhidan, M. Hamilton, and D. J. D’Souza, “Instructor perspec-
solving a problem unknown to MindReader, and the new CDG tives of multiple-choice questions in summative assessment for novice
is included in the knowledgebase after proper curation. programmers,” Computer Science Education, vol. 20, no. 3, pp. 229–
259, 2010.
V. S UMMARY AND F UTURE R ESEARCH [10] J. C. Gluz, F. Penteado, M. Mossmann, L. Gomes, and R. Vicari, “A
student model for teaching natural deduction based on a prover that
In this paper, we reported a late breaking result of a mimics student reasoning,” in ITS, Honolulu, HI, USA, June 5-9, 2014,
pp. 482–489.
research focusing on semantic understanding of student codes
in an online learning environment. We have demonstrated [11] S. Mehta, “Development and assessment of interactive web-based
problem solvers in reinforcing math and physics concepts,” in WWW
that CDG based matching code fragments have a higher WebNet, Orlando, Florida, October 23-27, 2001, p. 844.
likelihood of detecting semantic and functional equivalence [12] P. Návrat and J. Tvarozek, “Online programming exercises for sum-
of two programs. The process is complemented by a dataflow mative assessment in university courses,” in 15th CompSysTech, Ruse,
analysis and random testing regime to identify possible valid Bulgaria, June 27-28, 2014, pp. 341–348.
solutions and learn new rules. We have also demonstrated [13] J. Sorva and T. Sirkiä, “Embedded questions in ebooks on programming:
that detecting code clones and plagiarized codes based on useful for a) summative assessment, b) formative assessment, or c)
PDGs is fundamentally different from matching two codes something else?” in 15th Koli Calling Conference on Computing
Education Research, Koli, Finland, November 19-22, 2015, pp. 152–
functionally using CDGs. In CDGs we substitute equivalent 156.
nodes under the guidance of a template CDG, and concept [14] H. B. Hunt III, R. L. Constable, and S. Sahni, “On the computational
hierarchy to determine semantic similarity essential for grading complexity of program scheme equivalence,” SIAM J. Comput., vol. 9,
tasks of MOOC student assignments. It should be evident that no. 2, pp. 396–416, 1980.
summarization of concepts in the concept graphs allows for [15] T. Eiter, M. Fink, H. Tompits, and S. Woltran, “Complexity results
abstract algorithm development, and it should be possible to for checking equivalence of stratified logic programs,” in 20th IJCAI,
actually write codes in various languages as an inverse function Hyderabad, India, January 6-12, 2007, pp. 330–335.
and develop new languages such as Scratch. [16] W. Drabent and M. Milkowska, “Proving correctness and completeness
of normal programs - a declarative approach,” TPLP, vol. 5, no. 6, pp.
669–711, 2005.
R EFERENCES [17] H. Hungar, “Complexity of proving program correctness,” in TACS,
Sendai, Japan, September 24-27, 1991, pp. 459–474.
[1] C. Domı́nguez, A. J. Elizondo, A. Sánchez, J. M. Blanco, and J. Heras,
“A comparative analysis of the consistency and difference among [18] T. Tang, R. Smith, S. Rixner, and J. Warren, “Data-driven test case
online self-, peer-, external- and instructor-assessments: The competitive generation for automated programming assessment,” in ACM ITiCSE,
effect,” Computers in Human Behavior, vol. 60, pp. 112–120, 2016. Arequipa, Peru, July 9-13, 2016, pp. 260–265.
[2] R. L. Edwards, J. K. Stewart, and M. Ferati, “Assessing the effectiveness [19] T. Görg, “Interprocedural pdg-based code clone detection,”
of distributed pair programming for an online informatics curriculum,” Softwaretechnik-Trends, vol. 36, no. 2, 2016.
Inroads, vol. 1, no. 1, pp. 48–54, 2010. [20] S. Wagner, A. Abdulkhaleq, I. Bogicevic, J. Ostberg, and J. Ramadani,
[3] W. Farag, “Comparing achievement of intended learning outcomes in “How are functionally similar code clones syntactically different? an
online programming classes with blended offerings,” in ACM SIGITE, empirical study and a benchmark,” PeerJ Computer Science, vol. 2, p.
Calgary, AB, Canada, October 11 - 13, 2012, pp. 25–30. e49, 2016.
[4] J. Sharp, “Traditional, online, and flipped: A preliminary analysis of [21] F. Zhang, D. Wu, P. Liu, and S. Zhu, “Program logic based software
instructional approaches on student performance in a c# programming plagiarism detection,” in 25th IEEE ISSRE, Naples, Italy, November
course,” in 22nd AMCIS, San Diego, CA, USA, August 11-14, 2016. 3-6, 2014, pp. 66–77.
[22] S. Boumiza, D. Souilem, and A. Bekiarski, “Workflow approach to
[5] V. J. Martin, T. Pereira, S. Sridharan, and C. R. Rivero, “Automated
design automatic tutor in e-learning environment,” in CoDIT, Saint
personalized feedback in introductory java programming moocs,” in
Julian’s, Malta, April 6-8, 2016, pp. 263–268.
33rd IEEE ICDE, California, USA, April 19-22, 2017, to appear.
[23] C. Hulls, A. Neale, B. Komalo, V. Petrov, and D. J. Brush, “Interactive
[6] H. M. Jamil, “Smart assessment of and tutoring for computational think-
online tutorial assistance for a first programming course,” IEEE Trans.
ing MOOC assignments using MindReader,” Department of Computer
Education, vol. 48, no. 4, pp. 719–728, 2005.
Science, University of Idaho, Moscow, Idaho, Tech. Rep., February
2017.

263

You might also like