Index: /anr/anr.tex
===================================================================
--- /anr/anr.tex	(revision 11)
+++ /anr/anr.tex	(revision 12)
@@ -11,7 +11,19 @@
 \usepackage{graphicx}
 \usepackage{color}
-
-\definecolor{gris25}{gray}{0.75}
-\definecolor{gris75}{gray}{0.30}
+\usepackage{xspace}
+\usepackage{geometry}
+\geometry{verbose,a4paper,tmargin=3cm,bmargin=2cm,lmargin=2cm,rmargin=3cm}
+
+%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
+\def\xcoach{\texttt{xcoach}\xspace}
+\def\backbone{backbone\xspace}
+
+%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
+\def\anrdoc#1{}
+\let\pagefeed\relax
+\def\euro{\mbox{\raisebox{.25ex}{{\it =}}\hspace{-.5em}{\sf C}}}
+\definecolor{gris}{gray}{0.75}
+\definecolor{rouge}{rgb}{1.0,0.2,0.2}
+\def\mustbecompleted#1{}
 
 \title{%
@@ -21,9 +33,248 @@
 }
 
-%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
-
+%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
+% DEBUT CONFIG
+% Comment next marcro to suppress the printing of anr directives
+\def\anrdoc#1{\noindent\begin{scriptsize}\textcolor{red}{#1}\end{scriptsize}\ifhmode\par\fi}
+% Comment the next macro to suppress the pagefeed
+\let\pagefeed\newpage
+% Comment the next macro to suppress it
+\def\mustbecompleted#1{\textcolor{gris}{#1}}
+% FIN CONFIG
+%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
+
+%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
 \begin{document}
-\input{body.tex}
+%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
+
+%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
+% 1
+\section{Executive summary}
+\anrdoc{(2 pages maximum) Résumer la problématique que le projet propose
+de résoudre, et comment cet objectif sera poursuivi (quelle approche
+technique, etc.). \\
+Ce résumé devra démontrer l'originalité du projet notamment sur les points
+suivants:
+\begin{itemize}
+\item les objectifs globaux, les verrous scientifiques et techniques,
+\item le programme de travail,
+\item les retombées scientifiques, techniques et économiques.
+\end{itemize}
+Ces éléments peuvent être recopiés dans les champs «résumés scientifiques»
+du site de soumission.}
+\input{section-1}
+
+%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
+% 2
+\pagefeed\section{Context and relevance to the call }
+\anrdoc{(1 page maximum) Présentation générale du problème qu'il est
+proposé de traiter dans le projet et du cadre de travail
+(recherche fondamentale, industrielle ou développement expérimental).}
+\input{section-2}
+
+% 2.1
+\pagefeed\subsection{Context, economic and societal issues}
+\anrdoc{(2 pages maximum) Décrire le contexte économique, social, réglementaire
+dans lequel se situe le projet en présentant une analyse des enjeux sociaux,
+économiques, environnementaux, industriels. Donner si possible des arguments
+chiffrés, par exemple, pertinence et portée du projet par rapport à la
+demande économique (analyse du marché, analyse des tendances), analyse
+de la concurrence, indicateurs de réduction de coûts, perspectives de
+marchés (champs d'application, ...). Indicateurs des gains environnementaux,
+cycle de vie.}
+\input{section-2.1}
+
+% 2.2
+\pagefeed\subsection{Relevance of the proposal}
+\anrdoc{(2 pages maximum) Préciser :\begin{itemize}
+\item positionnement du projet par rapport au contexte développé précédemment:
+  vis- à-vis des projets et recherches concurrents, complémentaires ou
+  antérieurs, des brevets et standards...
+\item positionnement du projet par rapport aux axes thématiques de l'appel
+  à projets.
+\item positionnement du projet aux niveaux européen et international.
+\end{itemize}}
+\input{section-2.2}
+
+%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
+% 3
+\section{Scientific and technical Description}
+
+% 3.1
+\pagefeed\subsection{State of the Art}
+\anrdoc{(3 pages maximum) écrire le contexte et les enjeux scientifiques
+dans lequel se situe le projet en présentant un état de l'art national et
+international dressant l'état des connaissances sur le sujet. Faire
+apparaître d'éventuels résultats préliminaires. Inclure les références
+bibliographiques nécessaires en annexe 7.1.}
+\input{section-3.1}
+
+% 3.2
+\pagefeed\subsection{S \& T objectives, progress beyond the state of the art}
+\anrdoc{(2 pages maximum)
+Décrire les objectifs scientifiques/techniques du projet.\\
+Présenter l'avancée scientifique attendue. Préciser l'originalité et le
+caractère ambitieux du projet.\\
+Détailler les verrous scientifiques et techniques à lever par la
+réalisation du projet.\\
+Décrire éventuellement le ou les produits finaux développés à l'issue du
+projet  montrant le caractère innovant du projet.\\
+Présenter les résultats escomptés en proposant si possible des critères de
+réussite et d'évaluation adaptés au type de projet, permettant d'évaluer
+les résultats en fin de projet.\\
+Le cas échéant (programmes exigeant la pluridisciplinarité), démontrer
+l'articulation entre les disciplines scientifiques.}
+\input{section-3.2}
+
+%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
+% 4
+\section{Scientific and technical objectives / project description}
+
+% 4.1
+\pagefeed\subsection{Scientific Programme, Project structure}
+\anrdoc{(2 pages maximum)
+Présentez le programme scientifique et justifiez la décomposition en tâches
+du programme de travail en cohérence avec les objectifs poursuivis.\\
+Utilisez un diagramme pour présenter les liens entre les différentes tâches
+(organigramme technique)\\
+Les tâches représentent les grandes phases du projet. Elles sont en nombre
+limité.\\
+N'oubliez pas les activités et actions correspondant à la dissémination et
+à la valorisation.}
+\input{section-4.1}
+
+% 4.2
+\pagefeed\subsection{Project management}
+\anrdoc{(2 pages maximum)
+Préciser les aspects organisationnels du projet et les modalités de
+coordination (si possible individualisation d'une tâche coordination : cf.
+tâche 0 du document de soumission A).}
+\input{section-4.2}
+
+% 4.3
+\pagefeed\subsection{Description of the tasks}
+\anrdoc{(idéalement 1 ou 2 pages par tâche)
+Pour chaque tâche, décrire:\begin{itemize}
+\item les objectifs  de la tâche et éventuels indicateurs de succès,
+\item le responsable de la tâche et les partenaires impliqués (possibilité
+de l'indiquer sous forme graphique),
+\item le programme détaillé des travaux par tâche,
+\item les livrables de la tâche,
+\item les contributions des partenaires (le «qui fait quoi»),
+\item la description des méthodes et des choix techniques et de la manière
+dont les solutions seront apportées,
+\item les risques de la tâche et les solutions de repli envisagées.
+\end{itemize}}
+
+\subsubsection{Task 1}
+%\input{task-0}
+
+\subsubsection{Task 2}
+\input{task-1}
+
+
+\subsection{Tasks schedule, deliverables and milestones}
+\anrdoc{(3 pages maximum)\begin{itemize}
+\item Présenter sous forme graphique un échéancier des différentes tâches
+et leurs dépendances (diagramme de Gantt par exemple).
+\item Présenter un tableau synthétique  de l'ensemble des livrables du
+projet (numéro de tâche, date, intitulé, responsable).
+\item Préciser de façon synthétique les jalons scientifiques et/ou
+techniques, les principaux points de rendez-vous, les points bloquants ou
+aléas qui risquent de remettre en cause l'aboutissement du projet ainsi que
+les réunions de projet prévues.\end{itemize}}
+
+%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
+\section{Dissemination and exploitation of results.
+         Management of intellectual property}
+\anrdoc{(1 à 2 pages)\\
+Présenter les stratégies de valorisation des résultats :
+\begin{itemize}
+\item la communication scientifique;
+\item la communication auprès du grand public;
+\item la valorisation des résultats attendus;
+\item les retombées scientifiques, techniques, industrielles, économiques, 
+\item la place du projet dans la stratégie industrielle des entreprises partenaires du projet
+\item autres retombées (normalisation, information des pouvoirs publics, ...)
+\item les échéances et la nature des retombées technico- économiques attendues
+\item l'incidence éventuelle sur l'emploi, la création d'activités nouvelles.
+\end{itemize}
+Présenter les grandes lignes des modes de protection et d'exploitation des
+résultats\\
+Pour les projets partenariaux organismes de recherche/entreprises, les
+partenaires devront conclure, sous l'égide du coordinateur du projet, un
+accord de consortium dans un délai de un an si le projet est retenu pour
+financement.\\
+Pour les projets académiques, l'accord de consortium n'est pas obligatoire
+mais fortement conseillé.}
+
+%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
+\section{Consortium Description}
+
+\subsection{Partners description \& relevance, complementarity}
+\anrdoc{(maximum 0,5 page par partenaire) Décrire brièvement chaque
+partenaire et fournir ici les éléments permettant d'apprécier la
+qualification des partenaires dans le projet (le « pourquoi qui fait quoi
+»). Il peut s'agir de réalisations passées, d'indicateurs (publications,
+brevets), de l'intérêt du partenaire pour le projet.\\
+Montrer la complémentarité et la valeur ajoutée des coopérations entre les
+différents partenaires. L'interdisciplinarité et l'ouverture à diverses
+collaborations seront à justifier en accord avec les orientations du
+projet. (1 page maximum)}
+
+\subsection{Relevant experience of the project coordinator}
+\anrdoc{(0,5 page maximum) Fournir les éléments permettant de juger la
+capacité du coordinateur à coordonner le projet.}
+
+%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
+\section{Scientific justification for the mobilisation of the  resources}
+\anrdoc{On présentera ici la justification scientifique et technique des moyens
+demandés dans le document de soumission A par chaque partenaire et
+synthétisés à l'échelle du projet dans la fiche «Tableaux récapitulatifs»
+du document de soumission A.\\
+Chaque partenaire justifiera les moyens qu'il demande en distinguant les
+différents postes de dépenses.}
+
+\subsection{Partner 1: XXX}
+\anrdoc{\begin{itemize}
+\item Equipment: 1) Préciser la nature des équipements* et justifier le
+    choix des équipements. 2) Si nécessaire, préciser la part de financement
+    demandé sur le projet et si les achats envisagés doivent être complétés
+    par d'autres sources de financement. Si tel est le cas, indiquer le
+    montant et l'origine de ces financements complémentaires.
+    3) Attention: Un devis sera demandé si le projet est retenu pour
+    financement.
+\item Personnel costs
+    1) Le personnel non permanent (thèses, post- doctorants,CDD...)
+    financé sur le projet devra être justifié.
+    2) Fournir  les profils des postes à pourvoir pour les personnels à
+    recruter (1/2 page maximum par type de poste)
+    3) Pour les thèses, préciser si des demandes de bourse de thèse sont
+    prévues ou en cours, en préciser la nature et la part de financement
+    imputable au projet. 
+\item Subcontracting. Préciser: 1) la nature des prestations
+    2) le type de prestataire.
+\item Travel.  Préciser: 1) les missions liées aux travaux d'acquisition
+    sur le terrain (campagnes de mesures),
+    2) les missions relevant de colloques, congrès.
+\item Expenses for inward billing (Costs justified by internal procedures
+of invoicing). Préciser la nature des prestations
+\item Other working costs. Toute dépense significative relevant de ce poste
+devra être justifiée.
+\end{itemize}}
+
+\subsection{Partner 2: XXX}
+
+%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
+\section{Annexes}
+%\subsection{References}
+\anrdoc{Inclure la liste des références bibliographiques utilisées dans la
+partie «Etat de l'art» et les références bibliographiques des
+partenaires ayant trait au projet.}
+\bibliographystyle{plain}
+\bibliography{anr}
+\newpage
+
+%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
 \end{document}
-
-%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
+%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
Index: /anr/architecture-csg.fig
===================================================================
--- /anr/architecture-csg.fig	(revision 12)
+++ /anr/architecture-csg.fig	(revision 12)
@@ -0,0 +1,112 @@
+#FIG 3.2  Produced by xfig version 3.2.5-alpha5
+Landscape
+Center
+Metric
+Letter  
+100.00
+Single
+-2
+1200 2
+6 3892 1126 5532 2807
+5 1 0 2 0 7 50 -1 -1 0.000 0 1 1 0 4886.631 1711.404 4857 2042 5139 1927 5211 1641
+	1 1 2.00 57.15 114.30
+6 4397 1798 4861 2262
+1 4 0 2 0 7 50 -1 -1 0.000 1 0.0000 4629 2030 217 217 4412 2030 4846 2030
+4 0 0 50 -1 2 12 0.0000 4 135 225 4515 2099 T0\001
+-6
+6 4217 2372 5192 2732
+4 1 0 50 -1 2 16 0.0000 4 195 855 4704 2507 process\001
+4 1 0 50 -1 2 16 0.0000 4 195 975 4704 2732 network\001
+-6
+1 4 0 2 0 7 50 -1 -1 0.000 1 0.0000 4214 1427 217 217 3996 1427 4431 1427
+1 4 0 2 0 7 50 -1 -1 0.000 1 0.0000 5180 1439 217 217 4963 1439 5398 1439
+2 1 0 2 0 7 50 -1 -1 0.000 0 0 -1 1 0 2
+	1 1 2.00 57.15 114.30
+	 4997 1427 4425 1427
+2 1 1 2 0 7 50 -1 -1 4.000 0 0 -1 0 0 2
+	 3917 2332 5517 2332
+2 4 0 2 0 7 50 -1 -1 0.000 0 0 7 0 0 5
+	 5507 2792 3907 2792 3907 1141 5507 1141 5507 2792
+4 0 0 50 -1 2 12 0.0000 4 135 225 4139 1498 T1\001
+4 0 0 50 -1 2 12 0.0000 4 135 225 5068 1498 T2\001
+-6
+6 1359 1177 3159 2752
+2 4 0 2 0 7 50 -1 -1 6.000 0 0 7 0 0 5
+	 2934 2527 2934 1402 1584 1402 1584 2527 2934 2527
+4 1 0 50 -1 2 14 0.0000 4 165 705 2259 1852 initilal\001
+4 1 0 50 -1 2 14 0.0000 4 225 1185 2259 2152 application\001
+-6
+6 3945 135 5505 890
+4 1 0 50 -1 2 14 0.0000 4 165 1305 4725 300 architecture\001
+4 1 0 50 -1 2 14 0.0000 4 210 1110 4725 470 parameter\001
+4 1 0 50 -1 2 14 0.0000 4 120 135 4715 640 +\001
+4 1 0 50 -1 2 14 0.0000 4 240 1560 4725 830 hw/sw mapping\001
+-6
+6 7215 2150 8550 3080
+2 2 1 2 0 7 50 -1 -1 6.000 0 0 -1 0 0 5
+	 7230 2165 8535 2165 8535 3065 7230 3065 7230 2165
+4 1 0 50 -1 2 14 0.0000 4 165 855 7905 2435 harware\001
+4 1 0 50 -1 2 14 0.0000 4 210 1170 7905 2690 component\001
+4 1 0 50 -1 2 14 0.0000 4 165 900 7890 2960 libraries\001
+-6
+6 6255 3275 7905 4205
+2 2 1 2 0 7 50 -1 -1 6.000 0 0 -1 0 0 5
+	 6270 3290 7890 3290 7890 4190 6270 4190 6270 3290
+4 1 0 50 -1 2 14 0.0000 4 165 1305 7080 3560 architecture\001
+4 1 0 50 -1 2 14 0.0000 4 225 960 7080 3815 template\001
+4 1 0 50 -1 2 14 0.0000 4 225 735 7125 4085 library\001
+-6
+6 6525 1070 7895 1580
+2 2 1 2 0 7 50 -1 -1 6.000 0 0 -1 0 0 5
+	 6540 1085 7880 1085 7880 1565 6540 1565 6540 1085
+4 1 0 50 -1 2 16 0.0000 4 195 585 7220 1415 CSG\001
+-6
+6 5850 2150 6795 3095
+2 2 1 2 0 7 50 -1 -1 6.000 0 0 -1 0 0 5
+	 5865 2165 6780 2165 6780 3080 5865 3080 5865 2165
+4 1 0 50 -1 2 14 0.0000 4 165 600 6345 2435 micro\001
+4 1 0 50 -1 2 14 0.0000 4 165 705 6345 2690 kernel\001
+4 1 0 50 -1 2 14 0.0000 4 225 735 6330 2960 library\001
+-6
+2 1 0 2 0 7 50 -1 -1 6.000 0 0 -1 1 0 2
+	0 0 2.00 120.00 240.00
+	 2932 1922 3912 1922
+2 4 0 2 0 7 50 -1 -1 6.000 0 0 7 0 0 5
+	 10365 1595 10365 920 9240 920 9240 1595 10365 1595
+2 1 0 2 0 7 50 -1 -1 6.000 0 0 -1 1 0 2
+	0 0 2.00 120.00 240.00
+	 5625 455 6465 1250
+2 1 0 2 0 7 50 -1 -1 6.000 0 0 -1 1 0 2
+	0 0 2.00 120.00 240.00
+	 5505 1430 6435 1430
+2 1 0 2 0 7 50 -1 -1 6.000 0 0 -1 1 0 2
+	0 0 2.00 120.00 240.00
+	 7890 1220 9195 1220
+2 1 0 2 0 7 50 -1 -1 6.000 0 0 -1 1 0 2
+	0 0 2.00 120.00 240.00
+	 7890 1415 9210 1880
+2 2 1 2 0 7 50 -1 -1 6.000 0 0 -1 0 0 5
+	 6510 110 7875 110 7875 545 6510 545 6510 110
+2 1 0 2 0 7 50 -1 -1 6.000 0 0 -1 1 0 2
+	0 0 2.00 120.00 240.00
+	 7200 1085 7200 560
+2 1 0 2 0 7 50 -1 -1 6.000 0 0 -1 1 0 2
+	0 0 2.00 120.00 240.00
+	 6255 2105 6690 1595
+2 1 0 2 0 7 50 -1 -1 6.000 0 0 -1 1 0 2
+	0 0 2.00 120.00 240.00
+	 7695 2120 7575 1595
+2 1 0 2 0 7 50 -1 -1 6.000 0 0 -1 1 0 2
+	0 0 2.00 120.00 240.00
+	 7040 3260 7025 1565
+2 4 0 2 0 7 50 -1 -1 6.000 0 0 7 0 0 5
+	 10335 2290 10335 1840 9210 1840 9210 2290 10335 2290
+2 1 1 2 0 7 50 -1 -1 6.000 0 0 -1 1 0 2
+	0 0 2.00 120.00 240.00
+	 9767 2310 9767 3210
+4 1 0 50 -1 2 14 0.0000 4 165 1020 9810 1445 bitstream\001
+4 1 0 50 -1 2 14 0.0000 4 165 645 9810 1205 FPGA\001
+4 1 0 50 -1 2 16 0.0000 4 195 570 7185 425 HLS\001
+4 1 0 50 -1 2 16 0.0000 4 195 135 7425 905 2\001
+4 1 0 50 -1 2 14 0.0000 4 120 990 9810 3470 measures\001
+4 1 0 50 -1 2 14 0.0000 4 165 1005 9795 2110 simulator\001
Index: /anr/architecture-hls.fig
===================================================================
--- /anr/architecture-hls.fig	(revision 12)
+++ /anr/architecture-hls.fig	(revision 12)
@@ -0,0 +1,135 @@
+#FIG 3.2  Produced by xfig version 3.2.5-alpha5
+Landscape
+Center
+Metric
+Letter  
+100.00
+Single
+-2
+1200 2
+6 1044 114 2472 1542
+6 1044 114 2472 1542
+6 1270 430 2245 1225
+4 1 0 50 -1 2 16 0.0000 4 195 780 1757 625 task of\001
+4 1 0 50 -1 2 16 0.0000 4 195 975 1757 895 network\001
+4 1 0 50 -1 2 16 0.0000 4 195 855 1757 1165 process\001
+-6
+1 4 0 2 0 7 50 -1 -1 6.000 1 0.0000 1758 828 699 699 1061 791 2456 866
+-6
+-6
+6 6550 225 9715 1575
+6 8140 225 9715 1575
+6 8583 645 9318 1170
+4 1 0 50 -1 2 14 0.0000 4 225 735 8950 1110 library\001
+4 1 0 50 -1 2 14 0.0000 4 165 420 8950 810 cell\001
+-6
+2 4 1 2 0 7 50 -1 -1 6.000 0 0 7 0 0 5
+	 9490 1350 9490 450 8365 450 8365 1350 9490 1350
+-6
+2 2 1 2 0 7 50 -1 -1 6.000 0 0 -1 0 0 5
+	 7240 540 7690 540 7690 765 7240 765 7240 540
+2 2 1 2 0 7 50 -1 -1 6.000 0 0 -1 0 0 5
+	 7240 990 7690 990 7690 1215 7240 1215 7240 990
+2 1 0 2 0 7 50 -1 -1 6.000 0 0 -1 1 0 2
+	0 0 2.00 120.00 240.00
+	 8365 765 7690 630
+2 1 0 2 0 7 50 -1 -1 6.000 0 0 -1 1 0 2
+	0 0 2.00 120.00 240.00
+	 8320 990 7690 1125
+2 1 0 2 0 7 50 -1 -1 6.000 0 0 7 1 1 2
+	0 0 2.00 120.00 240.00
+	0 0 2.00 120.00 240.00
+	 6565 1080 7240 1080
+2 1 0 2 0 7 50 -1 -1 6.000 0 0 7 1 1 2
+	0 0 2.00 120.00 240.00
+	0 0 2.00 120.00 240.00
+	 6565 630 7240 630
+-6
+6 7150 1755 7690 2970
+2 2 1 2 0 7 50 -1 -1 6.000 0 0 -1 0 0 5
+	 7195 1800 7645 1800 7645 2025 7195 2025 7195 1800
+2 2 1 2 0 7 50 -1 -1 6.000 0 0 -1 0 0 5
+	 7195 2250 7645 2250 7645 2475 7195 2475 7195 2250
+2 2 1 2 0 7 50 -1 -1 6.000 0 0 -1 0 0 5
+	 7195 2700 7645 2700 7645 2925 7195 2925 7195 2700
+-6
+6 7150 -1170 7690 45
+2 2 1 2 0 7 50 -1 -1 6.000 0 0 -1 0 0 5
+	 7195 -1125 7645 -1125 7645 -900 7195 -900 7195 -1125
+2 2 1 2 0 7 50 -1 -1 6.000 0 0 -1 0 0 5
+	 7195 -675 7645 -675 7645 -450 7195 -450 7195 -675
+2 2 1 2 0 7 50 -1 -1 6.000 0 0 -1 0 0 5
+	 7195 -225 7645 -225 7645 0 7195 0 7195 -225
+-6
+2 2 0 2 0 7 50 -1 -1 6.000 0 0 7 0 0 5
+	 3150 225 4275 225 4275 900 3150 900 3150 225
+2 2 1 2 0 7 50 -1 -1 6.000 0 0 -1 0 0 5
+	 3150 900 4275 900 4275 1575 3150 1575 3150 900
+2 1 0 2 0 7 50 -1 -1 6.000 0 0 -1 1 0 2
+	0 0 2.00 120.00 240.00
+	 2472 872 3120 595
+2 1 0 2 0 7 50 -1 -1 6.000 0 0 -1 1 0 2
+	0 0 2.00 120.00 240.00
+	 4252 1262 4932 802
+2 1 1 2 0 7 50 -1 -1 6.000 0 0 7 0 0 2
+	 4947 900 5847 900
+2 4 1 2 0 7 50 -1 -1 6.000 0 0 7 0 0 5
+	 5847 1350 5847 450 4947 450 4947 1350 5847 1350
+2 1 2 2 0 7 50 -1 -1 4.500 0 0 -1 0 0 2
+	 4947 225 8275 225
+2 1 2 2 0 7 50 -1 -1 4.500 0 0 7 0 0 2
+	 4947 1575 9565 1577
+2 1 2 2 0 7 50 -1 -1 4.500 0 0 7 0 0 3
+	 8320 225 8320 0 9445 0
+2 1 0 2 0 7 50 -1 -1 6.000 0 0 -1 0 0 2
+	 6527 -1028 6527 2802
+2 1 0 2 0 7 50 -1 -1 6.000 0 0 7 1 1 2
+	0 0 2.00 120.00 240.00
+	0 0 2.00 120.00 240.00
+	 6527 2802 7197 2802
+2 1 0 2 0 7 50 -1 -1 6.000 0 0 7 1 1 2
+	0 0 2.00 120.00 240.00
+	0 0 2.00 120.00 240.00
+	 6527 2352 7197 2352
+2 1 0 2 0 7 50 -1 -1 6.000 0 0 7 1 1 2
+	0 0 2.00 120.00 240.00
+	0 0 2.00 120.00 240.00
+	 6527 1912 7197 1912
+2 1 0 2 0 7 50 -1 -1 6.000 0 0 7 1 1 2
+	0 0 2.00 120.00 240.00
+	0 0 2.00 120.00 240.00
+	 6527 -108 7197 -108
+2 1 0 2 0 7 50 -1 -1 6.000 0 0 7 1 1 2
+	0 0 2.00 120.00 240.00
+	0 0 2.00 120.00 240.00
+	 6527 -558 7187 -548
+2 1 0 2 0 7 50 -1 -1 6.000 0 0 7 1 1 2
+	0 0 2.00 120.00 240.00
+	0 0 2.00 120.00 240.00
+	 6527 -1008 7197 -1008
+2 1 0 2 0 7 50 -1 -1 6.000 0 0 7 1 1 2
+	0 0 2.00 120.00 240.00
+	0 0 2.00 120.00 240.00
+	 5847 900 6565 900
+2 1 0 2 0 7 50 -1 -1 6.000 0 0 -1 1 0 2
+	0 0 2.00 120.00 240.00
+	 7662 1932 8292 1932
+2 1 0 2 0 7 50 -1 -1 6.000 0 0 -1 1 0 2
+	0 0 2.00 120.00 240.00
+	 7642 2362 8272 2362
+2 1 0 2 0 7 50 -1 -1 6.000 0 0 -1 1 0 2
+	0 0 2.00 120.00 240.00
+	 7642 2812 8272 2812
+4 1 0 50 -1 2 14 0.0000 4 165 735 3690 1125 xcoach\001
+4 1 0 50 -1 2 14 0.0000 4 165 645 3690 1425 driver\001
+4 1 0 50 -1 2 14 0.0000 4 180 360 3690 630 gcc\001
+4 1 0 50 -1 2 14 0.0000 4 165 735 5397 735 xcoach\001
+4 1 0 50 -1 2 16 0.0000 4 150 165 5382 1215 +\001
+4 0 0 50 -1 2 14 0.0000 4 165 1005 5082 -480 front-end\001
+4 0 0 50 -1 2 14 0.0000 4 165 750 5127 2430 drivers\001
+4 0 0 50 -1 2 14 0.0000 4 165 975 8480 275 back-end\001
+4 0 0 50 -1 2 16 0.0000 4 225 1335 8352 3047 component\001
+4 0 0 50 -1 2 16 0.0000 4 195 825 8322 2862 VHDL\001
+4 0 0 50 -1 2 16 0.0000 4 255 1065 8292 2312 SystemC\001
+4 0 0 50 -1 2 16 0.0000 4 195 735 8312 2537 model\001
+4 0 0 50 -1 2 16 0.0000 4 195 990 8322 2032 C model\001
Index: /anr/architecture-hpc.fig
===================================================================
--- /anr/architecture-hpc.fig	(revision 12)
+++ /anr/architecture-hpc.fig	(revision 12)
@@ -0,0 +1,81 @@
+#FIG 3.2  Produced by xfig version 3.2.5-alpha5
+Landscape
+Center
+Metric
+Letter  
+100.00
+Single
+-2
+1200 2
+6 3375 1485 4725 2160
+4 1 0 50 -1 2 14 0.0000 4 165 1140 4050 1710 FPGA SoC\001
+4 1 0 50 -1 2 14 0.0000 4 210 450 4050 2010 part\001
+-6
+6 3825 -90 4275 585
+4 1 0 50 -1 2 14 0.0000 4 165 315 4050 135 PC\001
+4 1 0 50 -1 2 14 0.0000 4 210 450 4050 435 part\001
+-6
+6 5580 1530 7020 2070
+4 1 0 50 -1 2 14 0.0000 4 165 1395 6300 1710 comunication\001
+4 1 0 50 -1 2 14 0.0000 4 225 735 6300 2010 library\001
+-6
+6 675 225 2475 1800
+2 4 0 2 0 7 50 -1 -1 6.000 0 0 7 0 0 5
+	 2250 1575 2250 450 900 450 900 1575 2250 1575
+4 1 0 50 -1 2 14 0.0000 4 165 705 1575 900 initilal\001
+4 1 0 50 -1 2 14 0.0000 4 225 1185 1575 1200 application\001
+-6
+6 3600 630 4500 855
+2 1 0 2 0 7 50 -1 -1 6.000 0 0 -1 0 0 3
+	 3825 780 4050 705 4275 780
+-6
+6 3600 1170 4500 1395
+2 1 0 2 0 7 50 -1 -1 6.000 0 0 -1 0 0 3
+	 3825 1245 4050 1320 4275 1245
+-6
+2 2 0 2 0 7 50 -1 -1 0.000 0 0 7 0 0 5
+	 5625 -225 6975 -225 6975 675 5625 675 5625 -225
+2 2 1 2 0 7 50 -1 -1 6.000 0 0 -1 0 0 5
+	 5490 1350 7110 1350 7110 2250 5490 2250 5490 1350
+2 1 0 2 0 7 50 -1 -1 6.000 0 0 7 1 0 2
+	0 0 2.00 120.00 240.00
+	 2700 225 3375 225
+2 1 0 2 0 7 50 -1 -1 6.000 0 0 7 0 0 2
+	 2250 990 2722 982
+2 4 0 2 0 7 50 -1 -1 6.000 0 0 7 0 0 5
+	 4725 675 4725 -225 3375 -225 3375 675 4725 675
+2 4 0 2 0 7 50 -1 -1 6.000 0 0 7 0 0 5
+	 4725 2250 4725 1350 3375 1350 3375 2250 4725 2250
+2 1 0 2 0 7 50 -1 -1 6.000 0 0 -1 1 0 2
+	0 0 2.00 120.00 240.00
+	 2700 1800 3375 1800
+2 1 0 2 0 7 50 -1 -1 6.000 0 0 7 0 0 2
+	 2700 225 2700 1800
+2 1 0 2 0 7 50 -1 -1 6.000 0 0 -1 1 0 2
+	0 0 2.00 120.00 240.00
+	 4725 225 5625 225
+2 1 0 2 0 7 50 -1 -1 6.000 0 0 -1 1 0 2
+	0 0 2.00 120.00 240.00
+	 4725 1800 5625 450
+2 1 0 2 0 7 50 -1 -1 6.000 0 0 -1 1 0 2
+	0 0 2.00 120.00 240.00
+	 6300 1350 6300 675
+2 1 0 2 0 7 50 -1 -1 6.000 0 0 -1 1 0 2
+	0 0 2.00 120.00 240.00
+	 6975 225 7875 225
+2 4 0 2 0 7 50 -1 -1 6.000 0 0 7 0 0 5
+	 9000 450 9000 0 7875 0 7875 450 9000 450
+2 1 0 2 0 7 50 -1 -1 6.000 0 0 -1 0 0 2
+	 3825 810 3825 1215
+2 1 0 2 0 7 50 -1 -1 6.000 0 0 -1 0 0 2
+	 4275 765 4275 1215
+2 1 1 2 0 7 50 -1 -1 6.000 0 0 -1 1 0 2
+	0 0 2.00 120.00 240.00
+	 8432 470 8432 1370
+4 1 0 50 -1 2 14 0.0000 4 225 945 6300 360 compiler\001
+4 1 0 50 -1 2 14 0.0000 4 165 165 6300 90 C\001
+4 1 0 50 -1 2 14 0.0000 4 165 1155 8460 270 executable\001
+4 1 0 50 -1 2 16 0.0000 4 195 135 2925 1125 1\001
+4 1 0 50 -1 2 16 0.0000 4 195 135 4050 1125 2\001
+4 1 0 50 -1 2 14 0.0000 4 120 990 8475 1630 measures\001
+4 1 0 50 -1 2 16 0.0000 4 195 135 8655 900 3\001
Index: r/body.tex
===================================================================
--- /anr/body.tex	(revision 11)
+++ 	(revision )
@@ -1,795 +1,0 @@
-\section{Project context}
-\hspace{2cm}\begin{scriptsize}\begin{verbatim}
-% 1.	CONTEXTE ET POSITIONNEMENT DU PROJET
-% (1 page maximum) Prï¿œsentation gï¿œnï¿œrale du problï¿œme qu'il est proposï¿œ de traiter 
-% dans le projet et du cadre de travail (recherche fondamentale, industrielle ou 
-% dï¿œveloppement expï¿œrimental).
-\end{verbatim}
-\end{scriptsize}
-An embedded system is an application integrated into one or several chips
-in order to accelerate it or to embedd it into a small device such as a personal 
-digital assistant (PDA).
-This topic is investigated since 80s using Applications Specific Integrated Circuits (ASIC),
-Digital Signal Processing (DSP) and parallel computing on multiprocessor machines or networks.
-More recently, since end of 90s, other technologies appeared like Very Large Instruction Word (VLIW),
-Application Specific Instruction Processors (ASIP), System on Chip (SoC), 
-Multi-Processors SoC (MPSoC).
-\\
-During these last decades embedded system was reserved to major industrial companies targeting high volume market
-due to the design and fabrication costs.
-Nowadays Field Programmable Gate Arrays (FPGA), like Virtex5 from Xilinx and Stratix4 from Altera, 
-can implement a SoC with multiple processors and several coprocessors for less than 10K euros
-per item. In addition, High Level Synthesis (HLS) becomes more mature and allows to automate 
-design and to drastically decrease its cost in terms of man power. Thus, both FPGA and HLS 
-tend to spread over HPC for small companies targeting low volume markets.
-\par
-To get an efficient embedded system, designer has to take into account application characteristics when it 
-chooses one of the former technologies.
-This choice is not easy and in most cases designer has to try different technologies to retain the
-most adapted one.
-\\
-The first objective of COACH is to provide an open-source framework to design embedded system
-on FPGA device.
-COACH framework allows designer to explore various software/hardware partitions of the
-target application, to run timing and functional simulations and to generate automatically both
-the software and the synthesizable description of the hardware.
-The main topics of the project are:
-\begin{itemize} 
-\item
-Design space exploration: It consists in analysing the application runnig on FPGA, defining the target
-technology (SoC, MPSoC, ASIP, ...) and hardware/software partitioning of tasks depending on
-technology choice. This exploration is driven basically by throughput, latency and power consumption 
-criteria. 
-\item
-Micro-architectural exploration: When hardware components are required, the HLS tools of the framework
-generate them automatically. At this stage the framework provides various HLS tools allowing the
-micro-architectural space design exploration. The exploration criteria are also throughput, latency
-and power consumption.
-% FIXME
-%CA At this stage, preliminary source-level transformations will be
-%CA required to improve the efficiency of the target component.
-%CA COACH will also provide such facilities, such as automatic parallelization
-%CA and memory optimisation.
-\item
-Performance measurement: For each point of design space exploration, metrics of criteria are available
-such as throughput, latency, power consumption, area, memory allocation and data locality. 
-They are evaluated using virtual prototyping, estimation or analysing methodologies.
-\item
-Targeted hardware technology: The COACH description of system is independent of the FPGA family.
-Every point of the design exploration space can be implemented on any FPGA having the required resources.
-Basically, COACH handles both Altera and Xilinx FPGA families.
-\end{itemize}
-As an extension of embedded system design, COACH deals also with High Performance Computing (HPC).
-In HPC, the kind of targeted application is an existing one running on PC. COACH helps designer
-to accelerate it by migrating critical parts into a SoC implemented on a FPGA plugged to the PC bus.
-\par
-COACH is the result of the will of several laboratory to unify their know how and skills in the
-following domains: Operating system and hardware communication (TIMA, SITI), SoC and MPSoC (LIP6 and TIMA),
-ASIP (IRISA) and HLS (LIP6, Lab-STIC and LIP). The project objective is to integrate these various 
-domains into a unique free framework (licence ...) masking as much as possible these domains and its 
-different tools to the user.
-
-
-\subsection{Economical context and interest}
-\hspace{2cm}\begin{scriptsize}\begin{verbatim}
-% 1.1.	CONTEXTE ET ENJEUX ECONOMIQUES ET SOCIETAUX 
-% (2 pages maximum)
-% Dï¿œcrire le contexte ï¿œconomique, social, rï¿œglementaire. dans lequel se situe 
-% le projet en prï¿œsentant une analyse des enjeux sociaux, ï¿œconomiques, environnementaux, 
-% industriels. Donner si possible des arguments chiffrï¿œs, par exemple, pertinence et 
-% portï¿œe du projet par rapport ï¿œ la demande ï¿œconomique (analyse du marchï¿œ, analyse des 
-% tendances), analyse de la concurrence, indicateurs de rï¿œduction de coï¿œts, perspectives 
-% de marchï¿œs (champs d'application, .). Indicateurs des gains environnementaux, cycle 
-% de vie.
-\end{verbatim}
-\end{scriptsize}
-Microelectronic allows to integrate complicated functions into products, to increase their
-commercial attractivity and to improve their competitivity. Multimedia and communication
-sectors have taken advantage from microelectronics facilities thanks to developpment of
-design methodologies and tools for real time embedded systems. Many other sectors could
-benefit from microelectronics if these methologies and tools are adapted to their features.
-The Non Recurring Engineering (NRE) costs involded in designing and manufacturing an ASIC is 
-very high. It costs several milliars of euros for IC factory and several millions to fabricate
-a specific circuit for example a conservative estimate for a 65nm ASIC project is 10 million USD. 
-Consequently, it is generally unfeasible to design and fabricate ASICs in
-low volumes and ICs are designed to cover a broad applications spectrum at the cost of
-performance degradation.
-\\
-Today, FPGAs become important actors in the computational domain that was originally dominated
-by microprocessors and ASICs. Just like microprocessors FPGA based systems can be reprogrammed
-on a per-application basis. At the same time, FPGAs offer significant performance benefits over
-microprocessors implementation for a number of applications. Although these benefits are still
-generally an order of magnitude less than equivalent ASIC implementations, low costs 
-(500 euros to 10K euros), fast time to market and flexibility of FPGAs make them an attractive 
-choice for low-to-medium volume applications. 
-Since their introduction in the mid eighties, FPGAs evolved from a simple, 
-low-capacity gate array technology to devices (Altera STRATIX III, Xilinx Virtex V) that
-provide a mix of coarse-grained data path units, memory blocks, microprocessor cores, 
-on chip A/D conversion, and gate counts by millions. This high logic capacity allows to implement
-complex systems like multi-processors platform with application dedicated coprocessors. 
-Table~\ref{fpga_market} shows the estimation of FPGA worldwide market in the next years covering 
-various application domains. The ``high end'' lines concern only FPGA with high logic capacity able 
-to implement complex systems. 
-This market is in significant expansion and is estimated to 914\,M\$ in 2012.
-Using FPGA limits the NRE costs to design cost. This boosts the developpment of methodologies
-and tools to automize design and reduce its cost.
-\begin{table}\leavevmode\center
-\begin{tabular}{|l|l|l|l|}\hline
-Segment	        & 2010	& 2011	& 2012 \\\hline\hline
-Communications	& 1,867	& 1,946	& 2,096 \\
-High end	& 467	& 511	& 550 \\\hline
-Consumer	& 550	& 592	& 672 \\
-High end	& 53	& 62	& 75 \\\hline
-Automotive	& 243	& 286	& 358 \\
-High end	& -	& -	& - \\\hline
-Industrial	& 1,102	& 1,228	& 1,406 \\
-High end	& 177	& 188	& 207 \\\hline
-Military/Aereo	& 566	& 636	& 717 \\
-High end	& 56	& 65	& 82 \\\hline\hline
-Total FPGA/PLD	& 4,659	& 5,015	& 5,583 \\
-Total High-End  FPGA	& 753	& 826	& 914 \\\hline
-\end{tabular}
-\caption{\label{fga_market} Gartner estimation of worldwide FPGA/PLD consumption (Millions \$)}
-\end{table}
-\par
-Today, several companies (atipa, blue-arc, Bull, Chelsio, Convey, CRAY, DataDirect, DELL, hp, 
-Wild Systems, IBM, Intel, Microsoft, Myricom, NEC, nvidia etc) are making systems where demand 
-for very high performance (HPC) primes over other requirements. They tend to use the highest 
-performing devices like Multi-core CPUs, GPUs, large FPGAs, custom ICs and the most innovative 
-architectures and algorithms. Companies show up in different "traditional" applications and market 
-segments like computing clusters (ad-hoc), servers and storage, networking and Telecom, ASIC 
-emulation and prototyping, Mil/aero etc. HPC market size is estimated today by FPGA providers 
-to 214\,M\$. 
-This market is dominated by Multi-core CPUs and GPUs based solutions and the expansion 
-of FPGA-based solutions is limited by the flow automation. Nowadays, there are neither commercial 
-nor free tools covering the whole design process.
-For instance, with SOPC Builder from Altera, users can select and parameterize IP components 
-from an extensive drop-down list of communication, digital signal processor (DSP), microprocessor 
-and bus interface cores, as well as incorporate their own IP. Designers can then generate 
-a synthesized netlist, simulation test bench and custom software library that reflect the hardware 
-configuration.
-Nevertheless, SOPC Builder does not provide any facilities to synthesize coprocessors\emph{I
-(Steven) disagree : the C2H compiler bundled with SOPCBuilder does a pretty good job at this} and to
-simulate the platform at a high design level (system C). 
-In addition, SOPC Builder is proprietary and only works together with Altera's Quartus compilation
-tool to implement designs on Altera devices (Stratix, Arria, Cyclone).
-PICO [CITATION] and CATAPULT [CITATION] allow to synthesize coprocessors from a C++ description.
-Nevertheless, they can only deal with data dominated applications and they do not handle the
-platform level.
-The Xilinx System Generator for DSP [http://www.xilinx.com/tools/sysgen.htm] is a plug-in to 
-Simulink that enables designers to develop high-performance DSP systems for Xilinx FPGAs. 
-Designers can design and simulate a system using MATLAB and Simulink. The tool will then 
-automatically generate synthesizable Hardware Description Language (HDL) code mapped to Xilinx 
-pre-optimized algorithms. 
-However, this tool targets only DSP based algorithms.
-\\
-Consequently, designers developping an embedded system needs to master for example
-SoCLib for design exploration,
-SOPC Builde at the platform level, 
-PICO for synthesizing the data dominated coprocessors
-and Quartus for design implementation.
-This requires an important tools interfacing effort and makes the design process very complex 
-and achievable only by designers skilled in many domains.
-COACH project integrates all these tools in the same framework masking them to the user. 
-The objective is to allow \textbf{pure software} developpers to realize embedded systems.
-\par
-The combination of the framework dedicated to software developpers and FPGA target, allows to gain 
-market share over Multi-core CPUs and GPUs HPC based solutions. 
-Moreover, one can expect that small and even very small companies will be able to propose embedded 
-system and accelerating solutions for standard software applications with acceptable prices, thanks 
- to the elimination of huge hardware investment in opposite to ASIC based solution.
-\\
-This new market may explose like it was done by micro-computing in eighties. This success were due 
-to the low cost of first micro-computers (compared to main frame) and the advent of high level 
-programming languages that allow a high number of programmers to launch start-ups in software
-engineering.
-
-\subsection{Project position}
-\hspace{2cm}\begin{scriptsize}\begin{verbatim}
-% 1.2.	POSITIONNEMENT DU PROJET
-% (2 pages maximum)
-% Prï¿œciser :
-% -	positionnement du projet par rapport au contexte dï¿œveloppï¿œ prï¿œcï¿œdemment : 
-%   vis- ï¿œ-vis des projets et recherches concurrents, complï¿œmentaires ou antï¿œrieurs, 
-%   des brevets et standards.
-% - positionnement du projet par rapport aux axes thï¿œmatiques de l'appel ï¿œ projets.
-% - positionnement du projet aux niveaux europï¿œen et international.
-\end{verbatim}
-\end{scriptsize}
-The aim of this project is to propose an open-source framework for architecture synthesis
-targeting mainly field programmable gate array circuits (FPGA).
-\\% LIP6/TIMA
-To evaluate the different architectures, the project uses the prototyping platform
-of the SoCLIB ANR project (2006-2009).
-\\% IRISA
-The project will also borrow from the ROMA ANR project (2007-2009) and the ongoing 
-joint INRIA-STMicro Nano2012 project. In particular we will adapt existing pattern 
-extraction algorithms and datapath merging techniques to the synthesis of customized 
-ASIP processors.
-\\
-\textcolor{gris75}{Steven : Je propose de rajouter un lien avec le projet BioWic~:~on the HPC
-application side, we also hope to benefit from the experience in hardware acceleration of
-bioinformatic algorithms/workfows gathered by the CAIRN group in the context of the ANR
-BioWic project (2009-2011), so as to be able to validate the framework on 
-real-life HPC applications.}
-
-\par
-%%% 1 -- POUVEZ VOUS CHACUN AJOUTER SVP (SI POSSIBLE) UNE LIGNE
-%%% 1 -- REFERANT UN PROJET ANR OU EUROPEEN
-%%% 1 -- Projets europï¿œens ou ANR rï¿œutilisï¿œs ou continuï¿œs
-%%% 1 LIP6/TIMA/LAB-STIC OK
-Regarding the expertise in  High Level Synthesis (HLS), the project leverages on know-how acquired over 15 years
-with GAUT project developped in Lab-STIC laboratory and UGH project developped in LIP6 
-and TIMA laboratories. \\
-Regarding architecture synthesis skills, the project is based on a know-how acquired over 10 years
-with the COSY European project (1998-2000) and the DISYDENT project developped in LIP6.  \\
-%%% 1 IRISA OK
-Regarding Application Specific Instruction Processor (ASIP) design, the CAIRN group at INRIA Bretagne
-Atlantique benefits from several years of expertise in the domain of retargetable compiler (Armor/Calife
-since 1996, and the Gecos compilers since 2002).
-
-
-% LIP FIXME:UN:PEU:LONG ET HORS:SUJET
-%CA% The source-level transformations required by the HLS tools will be
-%CA% designed in the {\em polyhedral model}, a general framework
-%CA% initiated by Paul Feautrier 20 years ago.  The programs handled in
-%CA% the polyhedral model are such that loop iterators describe a
-%CA% polyhedron (hence the name). This includes most of the kernels used
-%CA% in embedded applications. This property allows to design precise
-%CA% analysis by means of integer programming techniques.
-%CA% %communaute active & internationale
-%CA% %transfert techno (Reservoir)
-%CA% The polyhedral community is very active, and the technological
-%CA% transfer has now started. Reservoir Labs inc., a company based in
-%CA% New-York, is currently integrating the last polyhedral developments
-%CA% in its commercial compiler.
-%CA% %transfert techno (gcc)
-%CA% Also, polyhedra are progressively migrating into the {\sc GNU Gcc}
-%CA% compiler, via {\sc Graphite}, a module initially developed by
-%CA% Sebastian Pop.
-%CA% %outils existants
-%CA% Several tools have been developed in the polyhedral community,
-%CA% such as {\sc Piplib} (parameter integer programming library), and
-%CA% {\sc Polylib}, a library providing set operations on polyhedra. Both
-%CA% tools are almost mandatory in polyhedral tools, and have reached
-%CA% a sufficient level of maturity to be considered as standard.
-%syntol & bee ???
-% FIN
-% and on more than 15 years of experience on parallel hardware generation
-% in the polyedral model in the CAIRN group (MMAlpha software
-% developped in the group since 1996).
-%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
-%%% 2 -- A COMPLETER (COURT)
-%%% 2 -- For polyedric transformation and memory optimization ... LIP 
-%%% 2 -- For ASIP IRISA
-%%% 2 -- For ... CITI
-%%% 2 -- For ... TIMA
-\par
-The SoCLIB ANR platform were developped by 11 laboratories and 6 companies. It allows to
-describe hardware architectures with shared memory space and to deploy software
-applications on them to evaluate their performance. 
-The heart of this platform is a library containing simulation models (in SystemC)
-of hardware IP cores such as processors, buses, networks, memories, IO controller.
-The platform provides also embedded operating systems and software/hardware
-communication components useful to implement applications quickly.
-However, the synthesisable description of IPs have to be provided by users. \\
-This project enhances SoCLib by providing synthesisable VHDL of standard IPs.
-In addition, HLS tools such as UGH and GAUT allow to get automatically a synthesisable 
-description of an IP (coprocessor) from a sequential algorithm.
-%\par
-%%% 2 IRISA ?
-%%% 2 ASIP tool such as ... IRISA
-%%% 2 ...
-%%% 2 Coach uses pattern extractions from ROMA
-%\par
-%%% 2 LIP ?
-\par
-The different points proposed in this project cover priorities defined by the commission 
-experts in the field of Information Technolgies Society (IST) for Embedded
-systems: <<Concepts, methods and tools for designing systems dealing with systems complexity
-and allowing to apply efficiently applications and various products on embedded platforms,
-considering resources constraints (delais, power, memory, etc.), security and quality
-services>>.
-\\
-Our team aims at covering all the steps of the design flow of architecture synthesis.
-Our project overcomes the complexity of using various synthesis tools and description 
-languages required today to design architectures.
-
-\section{Scientific and Technical Description}
-\subsection{State of the art}
-\hspace{2cm}\begin{scriptsize}\begin{verbatim}
-% 2.	DESCRIPTION SCIENTIFIQUE ET TECHNIQUE
-% 2.1.	ï¿œTAT DE L'ART
-% (3 pages maximum)
-% Dï¿œcrire le contexte et les enjeux scientifiques dans lequel se situe le projet 
-% en prï¿œsentant un ï¿œtat de l'art national et international dressant l'ï¿œtat des 
-% connaissances sur le sujet. Faire apparaï¿œtre d'ï¿œventuels rï¿œsultats prï¿œliminaires. 
-% Inclure les rï¿œfï¿œrences bibliographiques nï¿œcessaires en annexe 7.1.
-\end{verbatim}
-\end{scriptsize}
-Our project covers several critical domains in system design in order
-to achieve high performance computing. Starting from a high level description we aim 
-at generating automatically both hardware and software components of the system.
-
-\subsubsection{High Performance Computing}
-Accelerating high-performance computing (HPC) applications with field-programmable
-gate arrays (FPGAs) can potentially improve performance. 
-However, using FPGAs presents significant challenges [1].
-First, the operating frequency of an FPGA is low compared to a high-end microprocessor.
-Second, based on Amdahl law,  HPC/FPGA application performance is unusually sensitive 
-to the implementation quality [2].
-Finally, High-performance computing programmers are a highly sophisticated but scarce 
-resource. Such programmers are expected to readily use new technology but lack the time 
-to learn a completely new skill such as logic design [3]. 
-\\
-HPC/FPGA hardware is only now emerging and in early commercial stages, 
-but these techniques have not yet caught up. 
-Thus, much effort is required to develop design tools that translate high level
-language programs to FPGA configurations.
-
-\hspace{2cm}\begin{scriptsize}\begin{verbatim}
-[1] M.B. Gokhale et al., Promises and Pitfalls of Reconfigurable
-Supercomputing, Proc. 2006 Conf. Eng. of Reconfigurable
-Systems and Algorithms, CSREA Press, 2006, pp. 11-20;
-http://nis-www.lanl.gov/~maya/papers/ersa06_gokhale_paper.
-pdf.
-[2] D. Buell, Programming Reconfigurable Computers: Language
-Lessons Learned, keynote address, Reconfigurable Systems
-Summer Institute 2006, 12 July 2006; http://gladiator.
-ncsa.uiuc.edu/PDFs/rssi06/presentations/00_Duncan_Buell.pdf
-[3] T. Van Court et al., Achieving High Performance
-with FPGA-Based Computing, Computer, vol. 40, no. 3, 
-pp. 50-57, Mar. 2007, doi:10.1109/MC.2007.79
-\end{verbatim}
-\end{scriptsize}
-
-\subsubsection{System Synthesis}
-Today, several solutions for system design are proposed and commercialized. The most common are
-those provided by Altera and Xilinx to promote their FPGA devices.
-\\
-The Xilinx System Generator for DSP [http://www.xilinx.com/tools/sysgen.htm] is a plug-in to 
-Simulink that enables designers to develop high-performance DSP systems for Xilinx FPGAs. 
-Designers can design and simulate a system using MATLAB and Simulink. The tool will then 
-automatically generate synthesizable Hardware Description Language (HDL) code mapped to Xilinx 
-pre-optimized algorithms. 
-However, this tool targets only DSP based algorithms, Xilinx FPGAs and cannot handle complete
-SoC. Thus, it is not really a system synthesis tool.
-\\
-In the opposite, SOPC Builder [CITATION] allows to describe a system, to synthesis it, 
-to programm it into a target FPGA and to upload a software application. 
-% FIXME(C2H from Altera, marche vite mais ressource monstrueuse)
-Nevertheless, SOPC Builder does not provide any facilities to synthesize coprocessors.
-Users have to provide the synthesizable description with the feasible bus interface.
-\\
-In addition, Xilinx System Generator and SOPC are closed world since each one imposes
-their own IPs which are not interchangeable.
-We can conclude that the existing commercial or free tools does not coverthe whole system 
-synthesis process in a full automatic way. Moreover, they are bound to a particular device family
-and to IPs library.
-
-\subsubsection{High Level Synthesis}
-High Level Synthesis translates a sequential algorithmic description and a constraints set 
-(area, power, frequency, ...) to a micro-architecture at Register Transfer Level (RTL).
-Several academic and commercial tools are today available. 
-Most common tools are SPARK [HLS1], GAUT [HLS2], UGH [HLS3] in the academic world 
-and catapultC [HLS4], PICO [HLS5] and Cynthesizer [HLS6] in commercial world.
-Despite their maturity, their usage is restrained by:
-\begin{itemize}
-\item They do not respect accurately the frequency constraint when they target an FPGA device.
-Their error is about 10 percent. This is annoying when the generated component is integrated
-in a SoC since it will slow down the hole system.
-\item These tools take into account only one or few constraints simultaneously while realistic
-designs are multi-constrained. 
-Moreover, low power consumption constraint is mandatory for embedded systems. 
-However, it is not yet well handled by common synthesis tools.
-\item The parallelism is extracted from initial algorithm. To get more parallelism or to reduce
-the amout of required memory, the user must re-write it while there is techniques as polyedric 
-transformations to increase the intrinsec parallelism.
-\item Despite they have the same input language (C/C++), they are sensitive to the style in
-which the algorithm is written. Consequently, engineering work is required to swap from 
-a tool to another.
-\item The HLS tools are not integrated into an architecture and system exploration tool.
-Thus, a designer who needs to accelerate a software part of the system, must adapt it manually 
-to the HLS input dialect and performs engineering work to exploit the synthesis result 
-at the system level.
-\end{itemize}
-Regarding these limitations, it is necessary to create a new tool generation reducing the gap 
-between the specification of an heterogenous system and its hardware implementation.
-
-\hspace{2cm}\begin{scriptsize}\begin{verbatim}
-[HLS1] SPARK universite de californie San Diego
-[HLS2] GAUT UBS/Lab-STIC
-[HLS3] UGH
-[HLS4] catapultC Mentor
-[HLS5] PICO synfora
-[HLS6] Cynthesizer Forte design system 
-\end{verbatim}
-\end{scriptsize}
-
-\subsubsection{Application Specific Instruction Processors}
-
-ASIP (Application-Specific Instruction-Set Processor) are programmable processors in 
-which both the instruction and the micro architecture have been tailored to a given
- application domain (eg. video processing), or to a specific application. 
-This specialization usually offers a good compromise between performance (w.r.t a pure software
-implementation on an embeded CPU) and flexibility (w.r.t an application specific 
-hardware co-processor).
-In spite of their obvious advantages, using/designing ASIPs remains a difficult
-task, since it involves designing both a micro-architecture and a compiler for this
-architecture. Besides, to our knowledge, there is still no available open-source
-design flow\footnote{There are commercial tools such a } for ASIP design even if such a tool would
-be valuable in the context of a System Level design exploration tool.    
-
-In this context, ASIP design based on Instruction Set Extensions (ISEs) has 
-received a lot of interest [NIOSII,TENSILICA]%~\cite{NIOS2,ST70}, 
-as it makes micro architecture synthesis 
-more tractable \footnote{ISEs rely on a template micro-architecture in which 
-only a small fraction of the architecture has to be specialized}, and help ASIP
-designers to focus on compilers, for which there are still many open problems 
-[CODES04,FPGA08].
-This approach however has a strong weakness, since it also significantly reduces 
-opportunities for achieving good seedups (most speedup remain between 1.5x and 
-2.5x), since ISEs performance is generally tied down by I/O constraints as 
-they generally rely on the main CPU register file to access data.
-
-% (
-%automaticcaly extraction ISE candidates for application code \cite{CODES04}, 
-%performing efficient instruction selection and/or storage resource (register) 
-%allocation \cite{FPGA08}).  
- 
-
-To cope with this issue, recent approaches~[DAC09,DAC08]%\cite{DAC09,DAC08} 
-advocate the use of 
-micro-architectural ISE models in which the coupling between the processor micro-architecture
-and the ISE component is thightened up so as to allow the ISE to overcome the register 
-I/O limitations, however these approaches tackle the problem for a compiler/simulation 
-point of view and not address the problem of generating synthesizable representations for 
-these models. 
-
-We therefore strongly believe that there is a need for an open-framework which
-would allow researchers and system designers to :
-\begin{itemize}
-\item Explore the various level of interactions between the original CPU micro-architecure
-and its extension (for example throught a Domain Specific Language targeted at micro-architecture
-specification and synthesis).
-\item Retarget the compiler instruction-selection (or prototype nex passes) passes so as
-to be able to take advantage of this ISEs.
-\item Provide  a complete System-level Integration for using ASIP as SoC building blocks 
-(integration with application specific blocks, MPSoc, etc.)
-\end{itemize}
-
-\hspace{2cm}
-\begin{scriptsize}\begin{verbatim} 
-
-[CODES08] Theo Kluter, Philip Brisk, Paolo Ienne, and Edoardo Charbon, Speculative DMA for
-Architecturally Visible Storage in Instruction Set Extensions
-
-[DAC09] Theo Kluter, Philip Brisk, Paolo Ienne, Edoardo Charbon, Way Stealing: Cache-assisted
-Automatic Instruction Set Extensions.
-
-[CODES04] Pan Yu, Tulika Mitra, Scalable Custom Instructions Identification for
-Instruction Set Extensible Processors.
-
-[FPGA08] Quang Dinh, Deming Chen, Martin D. F. Wong, Efficient ASIP Design for Configurable
-Processors with Fine-Grained Resource Sharing.
-
-[NIOSII] Nios II Custom Instruction User Guide
-
-\end{verbatim}
-
-\end{scriptsize}
-%, either 
-%because the target architecture is proprietary, or because the compiler 
-%technology is closed/commercial.
-
-
-
-
-% We propose to explore how to tighten the coupling of the extensions and 
-% the underlyoing template micro-architecture.
-% *  Thightne Even if such 
-% an approach offers less flexiblity and forbids very tight coupling 
-% between the extensions and the template micro-architecture, it makes the 
-% design of the micro-architecture more tractable and amenable to a fully 
-% automated flow.
-% \\
-% \\
-% In the context of the COACH project, we propose to add to the 
-% infra-structure a design flow targeted to automatic instruction set 
-% extension for the MIPS-based CPU, which will come as a complement or an 
-% alternative to the other proposed approaches (hardware accelerator, 
-% multi processors).
-% 
-
-\subsubsection{Automatic Parallelization}
-\begin{Large}\begin{verbatim}
--- A COMPLETER LIP
-\end{verbatim}
-\end{Large}
-%CA%   Parallel machines are often difficult and painful to program
-%CA%   directly, and one would like the compiler to %do the job, that is to
-%CA%   turn automatically a sequential program into a parallel form. This
-%CA%   transformation is referred as {\em automatic parallelization}, and has
-%CA%   been widely addressed since the 70s. Automatic parallelization
-%CA%   relies on data dependences, which cannot be computed in general.%, as
-%CA%   %one cannot predict at compile time the variable values on a given
-%CA%   %execution point. 
-%CA%   This negative result led researchers to (i) find a
-%CA%   program model in which no approximation is needed (ie polyhedral
-%CA%   model), (ii) make conservative approximations (iii) remark that
-%CA%   variable values are known at runtime, and make the decisions during
-%CA%   program execution. The latter approach is obviously not suitable
-%CA%   there, as we target hardware generation. We will give there a short
-%CA%   history of the approaches that fall in the first category.
-%CA%
-%CA%%   In the real world, we deal with a limited amount of processors,
-%CA%%   and the communication between processors takes time, and is
-%CA%%   critical for performance. %Whenever we have synchronisation-free
-%CA%%   parallelism, like for embarrassingly parallel kernels, this is not an
-%CA%%   issue. But in case of pipelined parallelism, we need to reduce
-%CA%%   communications as much as possible. 
-%CA%%   So we also need to find parallelism toghether with a proper mapping
-%CA%%   of operations and data on physical processors.
-%CA%
-%CA%   As programs spend most of there time in loops, the community has
-%CA%   focused on loop transformations that reveal parallelism. 
-%CA%%unimodulaire
-%CA%   The first approaches worked on perfect loop nests, where the tree
-%CA%   formed by the nested loops is linear. In this program model, the
-%CA%   loops can be seen as a basis that drive the way the iteration
-%CA%   domain will be described. Hence, a first idea was to change this
-%CA%   basis such that one vector (one loop) at least is parallel. To ease
-%CA%   the code generation, the area of defined by the news vectors must
-%CA%   be a unit volume. %Otherwise, one would produce an homothetic
-%CA%%   expansion of the iteration domain, which will force to put modulos
-%CA%%   in the target code. 
-%CA%   For this reason, these transformations are called {\em unimodular
-%CA%   transformations}.
-%CA%%tiling
-%CA%   
-%CA%   The next approaches include {\em loop tiling}, a simple
-%CA%   partitioning of the iteration domain, whose initial purpose is to
-%CA%   execute every partition on a different processor. %In the same way,
-%CA%   The execution order is modified with a proper unimodular
-%CA%   transformation, then the tiles are obtained by cutting the
-%CA%   iteration domain with the hyperplanes directed by every vector of
-%CA%   the new (unimodular) basis, at regular intervals. When the tiling
-%CA%   hyperplanes are properly chosen, we can both improve data-locality
-%CA%   on every processor, and reduce the communication between two
-%CA%   different tiles (which will be mapped on processors). This last
-%CA%   property implying that one tend to find a degree of parallelism as
-%CA%   great as possible.
-%CA%
-%CA%%affine scheduling
-%CA%   The previous approaches were restricted to kernels with perfect
-%CA%   loop nests (linear loop tree), and unimodular transformations. The
-%CA%   last generation of approaches broke with these limitations. We now
-%CA%   choose a different basis for every assignment, without the
-%CA%   unimodularity restriction. A dual way to present the things is the
-%CA%   notion of {\em affine schedule}, introduced by Feautrier [part1],
-%CA%   that simply assigns an abstract execution date to every assignment
-%CA%   execution. As an assignment execution is exactly characterised by
-%CA%   the current value of the loops counters (iteration vector), the
-%CA%   affine schedule will be defined as an affine form of the iteration
-%CA%   vector (hence the 'affine'). The affine property allows to use
-%CA%   integer programming techniques to compute the schedule. With this
-%CA%   approach, additional techniques are required to allocate the
-%CA%   parallel operations and the data to processor in an efficient way
-%CA%   [griebl, feautrier].
-%CA%
-%CA%%modularity??
-%CA%%%    As loop nests are no longer perfect, we deal with (transformed)
-%CA%%%    iteration domains of different dimensions, which can possibly (and
-%CA%%%    certainly) overlap. At this point, a new code generation technique
-%CA%%%    was needed. The first attempt is due to Chamsky et al. [??], and
-%CA%%%    was improved by Quillere et al. [QRW]. The code is now implemented
-%CA%%%    in an efficient tool [cloog], that gave a new life to polyhedral
-%CA%%%    techniques.
-%CA%
-%CA%%pluto's tiling
-%CA%   The tiling techniques were extended to non-perfect loop nest with
-%CA%   {\em affine partitioning}. Affine partitioning is to affine
-%CA%   scheduling what (original) tiling was to unimodular
-%CA%   transformations. An affine partitioning assigns to every assignment
-%CA%   its coordinates in the basis defined by the normals to the tiling
-%CA%   hyperplanes. Recently, a way to compute efficient hyperplanes were
-%CA%   found [uday], with a good data locality, and communications
-%CA%   confined in a small neighborhood around every processor.
-%CA%
-%CA%\subsubsection{Source-level Memory Optimisation}
-%CA%  The HLS process allows to customise memory, which impacts on final
-%CA%  circuit size and power consumption. Though most HLS tools already
-%CA%  try to optimise memory usage, it is better to provide an independent
-%CA%  source-level pass, that could be reused for different tools and in
-%CA%  other contexts.
-%CA%
-%CA%  There exists many approaches to evaluate and reduce the memory
-%CA%  requirement of a program. The first approaches are concerned with
-%CA%  {\em memory size estimation}, which can be defined as the maximum
-%CA%  number of memory cells used at the same time [clauss,zhao]. These
-%CA%  approaches provide an estimation as a symbolic expression of program
-%CA%  parameters, which can be used further to guide loop optimisations.
-%CA%  However, no explicit way to reduce the memory size is given.  {\em
-%CA%  Intra-array reuse} approaches brake with this limitation, and
-%CA%  collapse the array cells which are not alive at the same time. The
-%CA%  collapse is done by means of a data layout transformation, specified
-%CA%  with a linear (modular) mapping.  The first approaches were
-%CA%  developed at IMEC [balasa,catthoor], and basically try to linearize
-%CA%  the arrays and fold them using a modulo operator. Then, Lefebvre et
-%CA%  al. propose a solution to fold independently the array dimensions
-%CA%  [lefebvre]. Finally, Darte et al. provide a general formalisation of
-%CA%  the problem, together with a solution that subsumes the previous
-%CA%  approaches [darte]. A first implementation was made with the tool
-%CA%  {\sc Bee}, but there are still many limitations.
-%CA%
-%CA%  \begin{itemize} 
-%CA%  \item The tool is restricted to regular programs, whereas more
-%CA%  general programs could be handled with a conservative array liveness
-%CA%  analysis.
-%CA%
-%CA%  \item Programs depending on parameters (inputs) are not handled,
-%CA%  which forbids to handle, for example, the body of tiled loops.
-%CA%
-%CA%  \item The new array layout can brake spatial locality, and then impact
-%CA%  performance and power consumption. One would like to get a mapping
-%CA%  that improve or, at least, preserve the spatial locality of the
-%CA%  program.
-%CA%
-%CA%  \item Finally, the final memory compaction strongly depends on the
-%CA%  program schedule, and is naturally hindered by the
-%CA%  parallelism. Consequently, there is a trade-off to find with
-%CA%  automatic parallelization. An ideal solution would be to reduce
-%CA%  memory usage, while preserving parallelism.  
-%CA%  \end{itemize}
-
-\subsubsection{Interfaces}
-\begin{Large}\begin{verbatim}
--- A COMPLETER INSA Etat de l'art
-\end{verbatim}
-\end{Large}
-%
-%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
-\subsection{Objectives and innovation aspects}
-\hspace{2cm}\begin{scriptsize}\begin{verbatim}
-% 2.2.	OBJECTIFS ET CARACTERE AMBITIEUX/NOVATEUR DU PROJET 
-% (2 pages maximum)
-% Dï¿œcrire les objectifs scientifiques/techniques du projet.
-% Prï¿œsenter l'avancï¿œe scientifique attendue. Prï¿œciser l'originalitï¿œ et le caractï¿œre 
-% ambitieux du projet.
-% Dï¿œtailler les verrous scientifiques et techniques ï¿œ lever par la rï¿œalisation du projet.
-% Dï¿œcrire ï¿œventuellement le ou les produits finaux dï¿œveloppï¿œs ï¿œ l'issue du projet  
-% montrant le caractï¿œre innovant du projet.
-% Prï¿œsenter les rï¿œsultats escomptï¿œs en proposant si possible des critï¿œres de rï¿œussite 
-% et d'ï¿œvaluation adaptï¿œs au type de projet, permettant d'ï¿œvaluer les rï¿œsultats en 
-% fin de projet.
-% Le cas ï¿œchï¿œant (programmes exigeant la pluridisciplinaritï¿œ), dï¿œmontrer l'articulation 
-% entre les disciplines scientifiques.
-\end{verbatim}
-\end{scriptsize}
-
-% les objectifs scientifiques/techniques du projet.
-The objectives of COACH project are to develop a complete framework to 
-HPC (accelerating solutions for existing software applications)
-and embedded applications (implementing an application on a low power standalone device).
-The design steps are presented figure 1.
-\begin{figure}[hbtp]\leavevmode\center
-  \includegraphics[width=.8\linewidth]{flow}
-  \caption{\label{coach-flow} COACH flow.}
-\end{figure}
-\begin{description}
-\item[HPC setup] Here the user splits the application into 2 parts: the host application
-which remains on PC and the SoC application which migrates on SoC. 
-The framework provides a simulation model allowing to evaluate the partitioning.
-\item[SoC design] In this phase, 
-The user can obtain simulators at different abstraction levels of the SoC by giving to COACH framework
-a SoC description.  
-This description consists of a process network corresponding to the SoC application, 
-an OS, an instance of a generic hardware platform
-and a mapping of processes on the platform components. The supported mapping are 
-software (the process runs on a SoC processor),
-XXXpeci (the process runs on a SoC processor enhanced with dedicated instructions),
-and hardware (the process runs into a coprocessor generated by HLS and plugged on the SoC bus).
-\item[Application compilation] Once SoC description is validated, COACH generates automatically
-an FPGA bitstream containing the hardware platform with SoC application software and 
-an executable containing the host application. The user can launch the application by
-loading the bitstream on FPGA and running the executable on PC.
-\end{description}
- 
-% l'avancee scientifique attendue. Preciser l'originalite et le caractere 
-% ambitieux du projet. 
-The main scientific contribution of the project is to unify various synthesis techniques
-(same input and output formats) allowing the user to swap without engineering effort
-from one to an other and even to chain them, for example, to run polyedric transformation 
-before synthesis.
-Another advantage of this framework is to provide different abstraction levels from
-a single description.
-Finally, this description is device family independent and its hardware implementation
-is automatically generated.
-
-% Detailler les verrous scientifiques et techniques a lever par la realisation du projet.
-System design is a very complicated task and in this project we try to simplify it
-as much as possible. For this purpose we have to deal with the following scientific
-and technological barriers.
-\begin{itemize}
-\item The main problem in HPC is the communication between the PC and the SoC.
-This problem has 2 aspects. The first one is the efficiency. The second is to 
-eliminate enginnering effort to implement it at different abstract levels.
-\item COACH design flow has a top-down approach. In the such case,
-the required performance of a coprocessor (run frequency, maximum cycles for
-a given computation, power consumption, etc) are imposed by the other system
-components. The challenge is to allow user to control accurately the synthesis
-process. For instance, the run frequency must not be a result of the RTL synthesis
-but a strict synthesis constraint.
-\item HLS tools are sensitive to the style in which the algorithm is written.
-In addition, they are are not integrated into an architecture and system 
-exploration tool.
-Consequently, engineering work is required to swap from a tool to another,
-to integrate the resulting simulation model to an architectural exploration tool 
-and to synthesize the generated RTL description.
-%CA Additionnal preprocessing, source-level transformations, are thus
-%CA required to improve the process.
-%CA Particularly, this includes parallelism exposure and efficient memory mapping.
-\item Most HLS tools translate a sequential algorithm into a coprocessor
-containing a single data-path and finite state machine (FSM). In this way,
-only the fine grained parallelism is exploited (ILP parallelism).
-The challenge is to identify the coarse grained parallelism and to generate,
-from a sequential algorithm, coprocessor containing multiple communicating
-tasks (data-paths and FSMs).
-\end{itemize}
-
-%Presenter les resultats escomptes en proposant si possible des criteres de reussite 
-%et d'evaluation adaptes au type de projet, permettant d'evaluer les resultats en 
-%fin de projet.
-The main result is the framework. It is composed concretely of: 
-2 HPC communication shemes with their implementation, 
-5 HLS tools (control dominated HLS, data dominated HLS, Coarse grained HLS, 
-Memory optimisation HLS and ASIP),
-3 systemC based virtual prototyping environment extended with synthesizable
-RTL IP cores (generic, ALTERA/NIOS/AVALON, XILINX/MICROBLAZE/OPB),
-one design space exploration tool,
-one operating system (OS).
-\\
-The framework fonctionality will be demonstrated with XXX-EXAMPLE1, XXX-EXAMPLE2
-and XXX-EXAMPLE3 on 4 archictures (generic/XILINX, generic/ALTERA,
-proprietary/XILINX, proprietary/ALTERA).
-
-%% \section{}
-%% %3.	PROGRAMME SCIENTIFIQUE ET TECHNIQUE, ORGANISATION DU PROJET
-%% \subsection{}
-%% %3.1.	PROGRAMME SCIENTIFIQUE ET STRUCTURATION DU PROJET 
-%% %(2 pages maximum)
-%% %Prï¿œsentez le programme scientifique et justifiez la dï¿œcomposition en tï¿œches du 
-%% %programme de travail en cohï¿œrence avec les objectifs poursuivis. 
-%% %Utilisez un diagramme pour prï¿œsenter les liens entre les diffï¿œrentes tï¿œches 
-%% %(organigramme technique)
-%% %Les tï¿œches reprï¿œsentent les grandes phases du projet. Elles sont en nombre limitï¿œ.
-%% %N'oubliez pas les activitï¿œs et actions correspondant ï¿œ la dissï¿œmination et ï¿œ la 
-%% %valorisation.
-%% 
-%% %METTRE UNE FIGURE ICI DECRIVANT LES TACHES ET LEURS INTERACTION (AVEC LE FLOT  
-%% %EN FILIGRANE ? )
-%% \subsection{}
-%% %3.2.	MANAGEMENT DU PROJET
-%% %(2 pages maximum)
-%% %Prï¿œciser les aspects organisationnels du projet et les modalitï¿œs de coordination 
-%% %(si possible individualisation d'une tï¿œche coordination : cf. tï¿œche 0 du document 
-%% %de soumission A).
-%% \subsection{}
-%% %3.3.	DESCRIPTION DES TRAVAUX PAR TACHE
-%% %(idï¿œalement 1 ou 2 pages par tï¿œche)
-%% %Pour chaque tï¿œche, dï¿œcrire : 
-%% %-	les objectifs  de la tï¿œche et ï¿œventuels indicateurs de succï¿œs,
-%% %-	le responsable de la tï¿œche et les partenaires impliquï¿œs (possibilitï¿œ de 
-%% %l'indiquer sous forme graphique),
-%% %-	le programme dï¿œtaillï¿œ des travaux par tï¿œche,
-%% %-	les livrables de la tï¿œche,
-%% %-	les contributions des partenaires (le " qui fait quoi "),
-%% %-	la description des mï¿œthodes et des choix techniques et de la maniï¿œre dont 
-%% %les solutions seront apportï¿œes,
-%% %-	les risques de la tï¿œche et les solutions de repli envisagï¿œes.
-
-
-
-
-
-
Index: r/coach_irisa.bib
===================================================================
--- /anr/coach_irisa.bib	(revision 11)
+++ 	(revision )
@@ -1,33 +1,0 @@
-@InProceedings{KluterCodes08,
-  author = 	 {{Theo Kluter and  Philip Brisk and  Paolo Ienne and  and Edoardo Charbon}},
-  title = 	 {{Speculative DMA for Architecturally Visible Storage in Instruction Set Extensions}},
-  booktitle = {ISSS/CODES},
-  year = 	 {2008},
-}
-
-@InProceedings{KluterDAC09,
-  author = 	 {{Theo Kluter and  Philip Brisk and  Paolo Ienne and  and Edoardo Charbon}},
-  title = 	 {{Way Stealing : Cache-assisted Automatic Instruction Set Extensions}},
-  booktitle = {Design Automation Conference (DAC)},
-  year = 	 {2009},
-}
-
-@InProceedings{YuCodes04,
-  author = 	 {{Pan Yu and Tulika Mitra}},
-  title = 	 {{Scalable Custom Instructions Identification for Instruction Set Extensible Processors}},
-  booktitle = {ISSS/CODES},
-  year = 	 {2004},
-}
-
-@InProceedings{Dinh08,
-  author = 	 {{Quang Dinh and Deming Chen and Martin D.~F.~Wong}},
-  title = 	 {{Efficient ASIP Design for Configurable Processors with Fine-Grained Resource Sharing}},
-  booktitle = {ACM Internatibnal Conference Field Programmable Gate Arrays (FPGA)},
-  year = 	 {2008},
-}
-
-@Misc{NIOS2UG,
-  title = 	 {{Nios II Custom Instruction User Guide, Altera Corp.}},
-  year = 	 {2008},
-  
-}
Index: /anr/dependence-dev.fig
===================================================================
--- /anr/dependence-dev.fig	(revision 12)
+++ /anr/dependence-dev.fig	(revision 12)
@@ -0,0 +1,63 @@
+#FIG 3.2  Produced by xfig version 3.2.5-alpha5
+Landscape
+Center
+Metric
+Letter  
+100.00
+Single
+-2
+1200 2
+6 4635 -1575 7560 -1170
+4 1 0 50 -1 2 14 0.0000 4 165 120 4725 -1170 0\001
+4 1 0 50 -1 2 14 0.0000 4 165 120 5175 -1170 6\001
+4 1 0 50 -1 2 14 0.0000 4 165 240 5625 -1170 12\001
+4 1 0 50 -1 2 14 0.0000 4 165 240 6075 -1170 18\001
+4 1 0 50 -1 2 14 0.0000 4 165 240 6525 -1170 24\001
+4 1 0 50 -1 2 14 0.0000 4 165 240 6975 -1170 30\001
+4 1 0 50 -1 2 14 0.0000 4 165 240 7425 -1170 36\001
+4 1 0 50 -1 2 14 0.0000 4 165 675 6075 -1395 month\001
+-6
+2 1 1 1 0 7 50 -1 -1 4.000 0 0 7 0 0 2
+	 4725 -1080 4725 360
+2 1 1 1 0 7 50 -1 -1 4.000 0 0 7 0 0 2
+	 5175 -1080 5175 360
+2 1 1 1 0 7 50 -1 -1 4.000 0 0 7 0 0 2
+	 5625 -1080 5625 360
+2 1 1 1 0 7 50 -1 -1 4.000 0 0 7 0 0 2
+	 6075 -1080 6075 360
+2 1 1 1 0 7 50 -1 -1 4.000 0 0 7 0 0 2
+	 6525 -1080 6525 360
+2 1 1 1 0 7 50 -1 -1 4.000 0 0 7 0 0 2
+	 6975 -1080 6975 360
+2 1 1 1 0 7 50 -1 -1 4.000 0 0 7 0 0 2
+	 7425 -1080 7425 360
+2 1 0 1 0 7 50 -1 -1 4.000 0 0 -1 1 1 2
+	0 0 1.00 60.00 120.00
+	0 0 1.00 60.00 120.00
+	 4725 -900 6525 -900
+2 1 0 1 0 7 50 -1 -1 4.000 0 0 -1 1 1 2
+	0 0 1.00 60.00 120.00
+	0 0 1.00 60.00 120.00
+	 5220 -675 7425 -675
+2 1 0 1 0 7 50 -1 -1 4.000 0 0 -1 1 1 2
+	0 0 1.00 60.00 120.00
+	0 0 1.00 60.00 120.00
+	 5220 -450 7380 -450
+2 1 0 1 0 7 50 -1 -1 4.000 0 0 -1 1 1 2
+	0 0 1.00 60.00 120.00
+	0 0 1.00 60.00 120.00
+	 5220 -225 7380 -225
+2 1 0 1 0 7 50 -1 -1 4.000 0 0 -1 1 1 2
+	0 0 1.00 60.00 120.00
+	0 0 1.00 60.00 120.00
+	 4725 0 7425 0
+2 1 0 1 0 7 50 -1 -1 4.000 0 0 -1 1 1 2
+	0 0 1.00 60.00 120.00
+	0 0 1.00 60.00 120.00
+	 5220 225 7425 225
+4 0 0 50 -1 2 14 0.0000 4 165 630 3960 -855 task 1\001
+4 0 0 50 -1 2 14 0.0000 4 165 630 3960 -630 task 2\001
+4 0 0 50 -1 2 14 0.0000 4 165 630 3960 -405 task 3\001
+4 0 0 50 -1 2 14 0.0000 4 165 630 3960 -180 task 4\001
+4 0 0 50 -1 2 14 0.0000 4 165 630 3960 45 task 5\001
+4 0 0 50 -1 2 14 0.0000 4 165 630 3960 270 task 6\001
Index: /anr/dependence-test.fig
===================================================================
--- /anr/dependence-test.fig	(revision 12)
+++ /anr/dependence-test.fig	(revision 12)
@@ -0,0 +1,71 @@
+#FIG 3.2  Produced by xfig version 3.2.5-alpha5
+Landscape
+Center
+Metric
+Letter  
+100.00
+Single
+-2
+1200 2
+6 4635 -1575 7560 -1170
+4 1 0 50 -1 2 14 0.0000 4 165 120 4725 -1170 0\001
+4 1 0 50 -1 2 14 0.0000 4 165 120 5175 -1170 6\001
+4 1 0 50 -1 2 14 0.0000 4 165 240 5625 -1170 12\001
+4 1 0 50 -1 2 14 0.0000 4 165 240 6075 -1170 18\001
+4 1 0 50 -1 2 14 0.0000 4 165 240 6525 -1170 24\001
+4 1 0 50 -1 2 14 0.0000 4 165 240 6975 -1170 30\001
+4 1 0 50 -1 2 14 0.0000 4 165 240 7425 -1170 36\001
+4 1 0 50 -1 2 14 0.0000 4 165 675 6075 -1395 month\001
+-6
+2 1 1 1 0 7 50 -1 -1 4.000 0 0 7 0 0 2
+	 4725 -1080 4725 360
+2 1 1 1 0 7 50 -1 -1 4.000 0 0 7 0 0 2
+	 5175 -1080 5175 360
+2 1 1 1 0 7 50 -1 -1 4.000 0 0 7 0 0 2
+	 5625 -1080 5625 360
+2 1 1 1 0 7 50 -1 -1 4.000 0 0 7 0 0 2
+	 6075 -1080 6075 360
+2 1 1 1 0 7 50 -1 -1 4.000 0 0 7 0 0 2
+	 6525 -1080 6525 360
+2 1 1 1 0 7 50 -1 -1 4.000 0 0 7 0 0 2
+	 6975 -1080 6975 360
+2 1 1 1 0 7 50 -1 -1 4.000 0 0 7 0 0 2
+	 7425 -1080 7425 360
+2 1 0 1 0 7 50 -1 -1 4.000 0 0 -1 1 1 2
+	0 0 1.00 60.00 120.00
+	0 0 1.00 60.00 120.00
+	 4725 -900 6525 -900
+2 1 0 1 0 7 50 -1 -1 4.000 0 0 -1 1 1 2
+	0 0 1.00 60.00 120.00
+	0 0 1.00 60.00 120.00
+	 5625 -450 7380 -450
+2 1 0 1 0 7 50 -1 -1 4.000 0 0 -1 1 1 2
+	0 0 1.00 60.00 120.00
+	0 0 1.00 60.00 120.00
+	 5625 -225 7380 -225
+2 1 0 1 0 7 50 -1 -1 4.000 0 0 -1 1 1 2
+	0 0 1.00 60.00 120.00
+	0 0 1.00 60.00 120.00
+	 6075 225 7425 225
+2 1 0 1 0 7 50 -1 -1 4.000 0 0 -1 1 1 2
+	0 0 1.00 60.00 120.00
+	0 0 1.00 60.00 120.00
+	 5220 -765 7425 -765
+2 1 0 1 0 7 50 -1 -1 4.000 0 0 -1 1 1 2
+	0 0 1.00 60.00 120.00
+	0 0 1.00 60.00 120.00
+	 6075 -630 7425 -630
+2 1 0 1 0 7 50 -1 -1 4.000 0 0 -1 1 1 2
+	0 0 1.00 60.00 120.00
+	0 0 1.00 60.00 120.00
+	 4725 -90 7425 -90
+2 1 0 1 0 7 50 -1 -1 4.000 0 0 -1 1 1 2
+	0 0 1.00 60.00 120.00
+	0 0 1.00 60.00 120.00
+	 6075 45 7425 45
+4 0 0 50 -1 2 14 0.0000 4 165 630 3960 -855 task 1\001
+4 0 0 50 -1 2 14 0.0000 4 165 630 3960 -630 task 2\001
+4 0 0 50 -1 2 14 0.0000 4 165 630 3960 -405 task 3\001
+4 0 0 50 -1 2 14 0.0000 4 165 630 3960 -180 task 4\001
+4 0 0 50 -1 2 14 0.0000 4 165 630 3960 45 task 5\001
+4 0 0 50 -1 2 14 0.0000 4 165 630 3960 270 task 6\001
Index: /anr/obsolete/anr.tex
===================================================================
--- /anr/obsolete/anr.tex	(revision 12)
+++ /anr/obsolete/anr.tex	(revision 12)
@@ -0,0 +1,29 @@
+\documentclass[12pt,a4paper]{article}
+
+\usepackage[french]{babel}
+%\usepackage[utf8x]{inputenc}
+\usepackage{times}
+\usepackage[T1]{fontenc}
+\usepackage{aeguill}
+\usepackage{verbatim}
+\usepackage{algorithm,algorithmic}
+\usepackage{xmpmulti}
+\usepackage{graphicx}
+\usepackage{color}
+
+\definecolor{gris25}{gray}{0.75}
+\definecolor{gris75}{gray}{0.30}
+
+\title{%
+\textbf{COACH:}
+\textbf{C}onception d'\textbf{A}rchitecture par
+\textbf{C}ompilation et synt\textbf{H}ï¿œse
+}
+
+%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
+
+\begin{document}
+\input{body.tex}
+\end{document}
+
+%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
Index: /anr/obsolete/body.tex
===================================================================
--- /anr/obsolete/body.tex	(revision 12)
+++ /anr/obsolete/body.tex	(revision 12)
@@ -0,0 +1,795 @@
+\section{Project context}
+\hspace{2cm}\begin{scriptsize}\begin{verbatim}
+% 1.	CONTEXTE ET POSITIONNEMENT DU PROJET
+% (1 page maximum) Prï¿œsentation gï¿œnï¿œrale du problï¿œme qu'il est proposï¿œ de traiter 
+% dans le projet et du cadre de travail (recherche fondamentale, industrielle ou 
+% dï¿œveloppement expï¿œrimental).
+\end{verbatim}
+\end{scriptsize}
+An embedded system is an application integrated into one or several chips
+in order to accelerate it or to embedd it into a small device such as a personal 
+digital assistant (PDA).
+This topic is investigated since 80s using Applications Specific Integrated Circuits (ASIC),
+Digital Signal Processing (DSP) and parallel computing on multiprocessor machines or networks.
+More recently, since end of 90s, other technologies appeared like Very Large Instruction Word (VLIW),
+Application Specific Instruction Processors (ASIP), System on Chip (SoC), 
+Multi-Processors SoC (MPSoC).
+\\
+During these last decades embedded system was reserved to major industrial companies targeting high volume market
+due to the design and fabrication costs.
+Nowadays Field Programmable Gate Arrays (FPGA), like Virtex5 from Xilinx and Stratix4 from Altera, 
+can implement a SoC with multiple processors and several coprocessors for less than 10K euros
+per item. In addition, High Level Synthesis (HLS) becomes more mature and allows to automate 
+design and to drastically decrease its cost in terms of man power. Thus, both FPGA and HLS 
+tend to spread over HPC for small companies targeting low volume markets.
+\par
+To get an efficient embedded system, designer has to take into account application characteristics when it 
+chooses one of the former technologies.
+This choice is not easy and in most cases designer has to try different technologies to retain the
+most adapted one.
+\\
+The first objective of COACH is to provide an open-source framework to design embedded system
+on FPGA device.
+COACH framework allows designer to explore various software/hardware partitions of the
+target application, to run timing and functional simulations and to generate automatically both
+the software and the synthesizable description of the hardware.
+The main topics of the project are:
+\begin{itemize} 
+\item
+Design space exploration: It consists in analysing the application runnig on FPGA, defining the target
+technology (SoC, MPSoC, ASIP, ...) and hardware/software partitioning of tasks depending on
+technology choice. This exploration is driven basically by throughput, latency and power consumption 
+criteria. 
+\item
+Micro-architectural exploration: When hardware components are required, the HLS tools of the framework
+generate them automatically. At this stage the framework provides various HLS tools allowing the
+micro-architectural space design exploration. The exploration criteria are also throughput, latency
+and power consumption.
+% FIXME
+%CA At this stage, preliminary source-level transformations will be
+%CA required to improve the efficiency of the target component.
+%CA COACH will also provide such facilities, such as automatic parallelization
+%CA and memory optimisation.
+\item
+Performance measurement: For each point of design space exploration, metrics of criteria are available
+such as throughput, latency, power consumption, area, memory allocation and data locality. 
+They are evaluated using virtual prototyping, estimation or analysing methodologies.
+\item
+Targeted hardware technology: The COACH description of system is independent of the FPGA family.
+Every point of the design exploration space can be implemented on any FPGA having the required resources.
+Basically, COACH handles both Altera and Xilinx FPGA families.
+\end{itemize}
+As an extension of embedded system design, COACH deals also with High Performance Computing (HPC).
+In HPC, the kind of targeted application is an existing one running on PC. COACH helps designer
+to accelerate it by migrating critical parts into a SoC implemented on a FPGA plugged to the PC bus.
+\par
+COACH is the result of the will of several laboratory to unify their know how and skills in the
+following domains: Operating system and hardware communication (TIMA, SITI), SoC and MPSoC (LIP6 and TIMA),
+ASIP (IRISA) and HLS (LIP6, Lab-STIC and LIP). The project objective is to integrate these various 
+domains into a unique free framework (licence ...) masking as much as possible these domains and its 
+different tools to the user.
+
+
+\subsection{Economical context and interest}
+\hspace{2cm}\begin{scriptsize}\begin{verbatim}
+% 1.1.	CONTEXTE ET ENJEUX ECONOMIQUES ET SOCIETAUX 
+% (2 pages maximum)
+% Dï¿œcrire le contexte ï¿œconomique, social, rï¿œglementaire. dans lequel se situe 
+% le projet en prï¿œsentant une analyse des enjeux sociaux, ï¿œconomiques, environnementaux, 
+% industriels. Donner si possible des arguments chiffrï¿œs, par exemple, pertinence et 
+% portï¿œe du projet par rapport ï¿œ la demande ï¿œconomique (analyse du marchï¿œ, analyse des 
+% tendances), analyse de la concurrence, indicateurs de rï¿œduction de coï¿œts, perspectives 
+% de marchï¿œs (champs d'application, .). Indicateurs des gains environnementaux, cycle 
+% de vie.
+\end{verbatim}
+\end{scriptsize}
+Microelectronic allows to integrate complicated functions into products, to increase their
+commercial attractivity and to improve their competitivity. Multimedia and communication
+sectors have taken advantage from microelectronics facilities thanks to developpment of
+design methodologies and tools for real time embedded systems. Many other sectors could
+benefit from microelectronics if these methologies and tools are adapted to their features.
+The Non Recurring Engineering (NRE) costs involded in designing and manufacturing an ASIC is 
+very high. It costs several milliars of euros for IC factory and several millions to fabricate
+a specific circuit for example a conservative estimate for a 65nm ASIC project is 10 million USD. 
+Consequently, it is generally unfeasible to design and fabricate ASICs in
+low volumes and ICs are designed to cover a broad applications spectrum at the cost of
+performance degradation.
+\\
+Today, FPGAs become important actors in the computational domain that was originally dominated
+by microprocessors and ASICs. Just like microprocessors FPGA based systems can be reprogrammed
+on a per-application basis. At the same time, FPGAs offer significant performance benefits over
+microprocessors implementation for a number of applications. Although these benefits are still
+generally an order of magnitude less than equivalent ASIC implementations, low costs 
+(500 euros to 10K euros), fast time to market and flexibility of FPGAs make them an attractive 
+choice for low-to-medium volume applications. 
+Since their introduction in the mid eighties, FPGAs evolved from a simple, 
+low-capacity gate array technology to devices (Altera STRATIX III, Xilinx Virtex V) that
+provide a mix of coarse-grained data path units, memory blocks, microprocessor cores, 
+on chip A/D conversion, and gate counts by millions. This high logic capacity allows to implement
+complex systems like multi-processors platform with application dedicated coprocessors. 
+Table~\ref{fpga_market} shows the estimation of FPGA worldwide market in the next years covering 
+various application domains. The ``high end'' lines concern only FPGA with high logic capacity able 
+to implement complex systems. 
+This market is in significant expansion and is estimated to 914\,M\$ in 2012.
+Using FPGA limits the NRE costs to design cost. This boosts the developpment of methodologies
+and tools to automize design and reduce its cost.
+\begin{table}\leavevmode\center
+\begin{tabular}{|l|l|l|l|}\hline
+Segment	        & 2010	& 2011	& 2012 \\\hline\hline
+Communications	& 1,867	& 1,946	& 2,096 \\
+High end	& 467	& 511	& 550 \\\hline
+Consumer	& 550	& 592	& 672 \\
+High end	& 53	& 62	& 75 \\\hline
+Automotive	& 243	& 286	& 358 \\
+High end	& -	& -	& - \\\hline
+Industrial	& 1,102	& 1,228	& 1,406 \\
+High end	& 177	& 188	& 207 \\\hline
+Military/Aereo	& 566	& 636	& 717 \\
+High end	& 56	& 65	& 82 \\\hline\hline
+Total FPGA/PLD	& 4,659	& 5,015	& 5,583 \\
+Total High-End  FPGA	& 753	& 826	& 914 \\\hline
+\end{tabular}
+\caption{\label{fga_market} Gartner estimation of worldwide FPGA/PLD consumption (Millions \$)}
+\end{table}
+\par
+Today, several companies (atipa, blue-arc, Bull, Chelsio, Convey, CRAY, DataDirect, DELL, hp, 
+Wild Systems, IBM, Intel, Microsoft, Myricom, NEC, nvidia etc) are making systems where demand 
+for very high performance (HPC) primes over other requirements. They tend to use the highest 
+performing devices like Multi-core CPUs, GPUs, large FPGAs, custom ICs and the most innovative 
+architectures and algorithms. Companies show up in different "traditional" applications and market 
+segments like computing clusters (ad-hoc), servers and storage, networking and Telecom, ASIC 
+emulation and prototyping, Mil/aero etc. HPC market size is estimated today by FPGA providers 
+to 214\,M\$. 
+This market is dominated by Multi-core CPUs and GPUs based solutions and the expansion 
+of FPGA-based solutions is limited by the flow automation. Nowadays, there are neither commercial 
+nor free tools covering the whole design process.
+For instance, with SOPC Builder from Altera, users can select and parameterize IP components 
+from an extensive drop-down list of communication, digital signal processor (DSP), microprocessor 
+and bus interface cores, as well as incorporate their own IP. Designers can then generate 
+a synthesized netlist, simulation test bench and custom software library that reflect the hardware 
+configuration.
+Nevertheless, SOPC Builder does not provide any facilities to synthesize coprocessors\emph{I
+(Steven) disagree : the C2H compiler bundled with SOPCBuilder does a pretty good job at this} and to
+simulate the platform at a high design level (system C). 
+In addition, SOPC Builder is proprietary and only works together with Altera's Quartus compilation
+tool to implement designs on Altera devices (Stratix, Arria, Cyclone).
+PICO [CITATION] and CATAPULT [CITATION] allow to synthesize coprocessors from a C++ description.
+Nevertheless, they can only deal with data dominated applications and they do not handle the
+platform level.
+The Xilinx System Generator for DSP [http://www.xilinx.com/tools/sysgen.htm] is a plug-in to 
+Simulink that enables designers to develop high-performance DSP systems for Xilinx FPGAs. 
+Designers can design and simulate a system using MATLAB and Simulink. The tool will then 
+automatically generate synthesizable Hardware Description Language (HDL) code mapped to Xilinx 
+pre-optimized algorithms. 
+However, this tool targets only DSP based algorithms.
+\\
+Consequently, designers developping an embedded system needs to master for example
+SoCLib for design exploration,
+SOPC Builde at the platform level, 
+PICO for synthesizing the data dominated coprocessors
+and Quartus for design implementation.
+This requires an important tools interfacing effort and makes the design process very complex 
+and achievable only by designers skilled in many domains.
+COACH project integrates all these tools in the same framework masking them to the user. 
+The objective is to allow \textbf{pure software} developpers to realize embedded systems.
+\par
+The combination of the framework dedicated to software developpers and FPGA target, allows to gain 
+market share over Multi-core CPUs and GPUs HPC based solutions. 
+Moreover, one can expect that small and even very small companies will be able to propose embedded 
+system and accelerating solutions for standard software applications with acceptable prices, thanks 
+ to the elimination of huge hardware investment in opposite to ASIC based solution.
+\\
+This new market may explose like it was done by micro-computing in eighties. This success were due 
+to the low cost of first micro-computers (compared to main frame) and the advent of high level 
+programming languages that allow a high number of programmers to launch start-ups in software
+engineering.
+
+\subsection{Project position}
+\hspace{2cm}\begin{scriptsize}\begin{verbatim}
+% 1.2.	POSITIONNEMENT DU PROJET
+% (2 pages maximum)
+% Prï¿œciser :
+% -	positionnement du projet par rapport au contexte dï¿œveloppï¿œ prï¿œcï¿œdemment : 
+%   vis- ï¿œ-vis des projets et recherches concurrents, complï¿œmentaires ou antï¿œrieurs, 
+%   des brevets et standards.
+% - positionnement du projet par rapport aux axes thï¿œmatiques de l'appel ï¿œ projets.
+% - positionnement du projet aux niveaux europï¿œen et international.
+\end{verbatim}
+\end{scriptsize}
+The aim of this project is to propose an open-source framework for architecture synthesis
+targeting mainly field programmable gate array circuits (FPGA).
+\\% LIP6/TIMA
+To evaluate the different architectures, the project uses the prototyping platform
+of the SoCLIB ANR project (2006-2009).
+\\% IRISA
+The project will also borrow from the ROMA ANR project (2007-2009) and the ongoing 
+joint INRIA-STMicro Nano2012 project. In particular we will adapt existing pattern 
+extraction algorithms and datapath merging techniques to the synthesis of customized 
+ASIP processors.
+\\
+\textcolor{gris75}{Steven : Je propose de rajouter un lien avec le projet BioWic~:~on the HPC
+application side, we also hope to benefit from the experience in hardware acceleration of
+bioinformatic algorithms/workfows gathered by the CAIRN group in the context of the ANR
+BioWic project (2009-2011), so as to be able to validate the framework on 
+real-life HPC applications.}
+
+\par
+%%% 1 -- POUVEZ VOUS CHACUN AJOUTER SVP (SI POSSIBLE) UNE LIGNE
+%%% 1 -- REFERANT UN PROJET ANR OU EUROPEEN
+%%% 1 -- Projets europï¿œens ou ANR rï¿œutilisï¿œs ou continuï¿œs
+%%% 1 LIP6/TIMA/LAB-STIC OK
+Regarding the expertise in  High Level Synthesis (HLS), the project leverages on know-how acquired over 15 years
+with GAUT project developped in Lab-STIC laboratory and UGH project developped in LIP6 
+and TIMA laboratories. \\
+Regarding architecture synthesis skills, the project is based on a know-how acquired over 10 years
+with the COSY European project (1998-2000) and the DISYDENT project developped in LIP6.  \\
+%%% 1 IRISA OK
+Regarding Application Specific Instruction Processor (ASIP) design, the CAIRN group at INRIA Bretagne
+Atlantique benefits from several years of expertise in the domain of retargetable compiler (Armor/Calife
+since 1996, and the Gecos compilers since 2002).
+
+
+% LIP FIXME:UN:PEU:LONG ET HORS:SUJET
+%CA% The source-level transformations required by the HLS tools will be
+%CA% designed in the {\em polyhedral model}, a general framework
+%CA% initiated by Paul Feautrier 20 years ago.  The programs handled in
+%CA% the polyhedral model are such that loop iterators describe a
+%CA% polyhedron (hence the name). This includes most of the kernels used
+%CA% in embedded applications. This property allows to design precise
+%CA% analysis by means of integer programming techniques.
+%CA% %communaute active & internationale
+%CA% %transfert techno (Reservoir)
+%CA% The polyhedral community is very active, and the technological
+%CA% transfer has now started. Reservoir Labs inc., a company based in
+%CA% New-York, is currently integrating the last polyhedral developments
+%CA% in its commercial compiler.
+%CA% %transfert techno (gcc)
+%CA% Also, polyhedra are progressively migrating into the {\sc GNU Gcc}
+%CA% compiler, via {\sc Graphite}, a module initially developed by
+%CA% Sebastian Pop.
+%CA% %outils existants
+%CA% Several tools have been developed in the polyhedral community,
+%CA% such as {\sc Piplib} (parameter integer programming library), and
+%CA% {\sc Polylib}, a library providing set operations on polyhedra. Both
+%CA% tools are almost mandatory in polyhedral tools, and have reached
+%CA% a sufficient level of maturity to be considered as standard.
+%syntol & bee ???
+% FIN
+% and on more than 15 years of experience on parallel hardware generation
+% in the polyedral model in the CAIRN group (MMAlpha software
+% developped in the group since 1996).
+%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
+%%% 2 -- A COMPLETER (COURT)
+%%% 2 -- For polyedric transformation and memory optimization ... LIP 
+%%% 2 -- For ASIP IRISA
+%%% 2 -- For ... CITI
+%%% 2 -- For ... TIMA
+\par
+The SoCLIB ANR platform were developped by 11 laboratories and 6 companies. It allows to
+describe hardware architectures with shared memory space and to deploy software
+applications on them to evaluate their performance. 
+The heart of this platform is a library containing simulation models (in SystemC)
+of hardware IP cores such as processors, buses, networks, memories, IO controller.
+The platform provides also embedded operating systems and software/hardware
+communication components useful to implement applications quickly.
+However, the synthesisable description of IPs have to be provided by users. \\
+This project enhances SoCLib by providing synthesisable VHDL of standard IPs.
+In addition, HLS tools such as UGH and GAUT allow to get automatically a synthesisable 
+description of an IP (coprocessor) from a sequential algorithm.
+%\par
+%%% 2 IRISA ?
+%%% 2 ASIP tool such as ... IRISA
+%%% 2 ...
+%%% 2 Coach uses pattern extractions from ROMA
+%\par
+%%% 2 LIP ?
+\par
+The different points proposed in this project cover priorities defined by the commission 
+experts in the field of Information Technolgies Society (IST) for Embedded
+systems: <<Concepts, methods and tools for designing systems dealing with systems complexity
+and allowing to apply efficiently applications and various products on embedded platforms,
+considering resources constraints (delais, power, memory, etc.), security and quality
+services>>.
+\\
+Our team aims at covering all the steps of the design flow of architecture synthesis.
+Our project overcomes the complexity of using various synthesis tools and description 
+languages required today to design architectures.
+
+\section{Scientific and Technical Description}
+\subsection{State of the art}
+\hspace{2cm}\begin{scriptsize}\begin{verbatim}
+% 2.	DESCRIPTION SCIENTIFIQUE ET TECHNIQUE
+% 2.1.	ï¿œTAT DE L'ART
+% (3 pages maximum)
+% Dï¿œcrire le contexte et les enjeux scientifiques dans lequel se situe le projet 
+% en prï¿œsentant un ï¿œtat de l'art national et international dressant l'ï¿œtat des 
+% connaissances sur le sujet. Faire apparaï¿œtre d'ï¿œventuels rï¿œsultats prï¿œliminaires. 
+% Inclure les rï¿œfï¿œrences bibliographiques nï¿œcessaires en annexe 7.1.
+\end{verbatim}
+\end{scriptsize}
+Our project covers several critical domains in system design in order
+to achieve high performance computing. Starting from a high level description we aim 
+at generating automatically both hardware and software components of the system.
+
+\subsubsection{High Performance Computing}
+Accelerating high-performance computing (HPC) applications with field-programmable
+gate arrays (FPGAs) can potentially improve performance. 
+However, using FPGAs presents significant challenges [1].
+First, the operating frequency of an FPGA is low compared to a high-end microprocessor.
+Second, based on Amdahl law,  HPC/FPGA application performance is unusually sensitive 
+to the implementation quality [2].
+Finally, High-performance computing programmers are a highly sophisticated but scarce 
+resource. Such programmers are expected to readily use new technology but lack the time 
+to learn a completely new skill such as logic design [3]. 
+\\
+HPC/FPGA hardware is only now emerging and in early commercial stages, 
+but these techniques have not yet caught up. 
+Thus, much effort is required to develop design tools that translate high level
+language programs to FPGA configurations.
+
+\hspace{2cm}\begin{scriptsize}\begin{verbatim}
+[1] M.B. Gokhale et al., Promises and Pitfalls of Reconfigurable
+Supercomputing, Proc. 2006 Conf. Eng. of Reconfigurable
+Systems and Algorithms, CSREA Press, 2006, pp. 11-20;
+http://nis-www.lanl.gov/~maya/papers/ersa06_gokhale_paper.
+pdf.
+[2] D. Buell, Programming Reconfigurable Computers: Language
+Lessons Learned, keynote address, Reconfigurable Systems
+Summer Institute 2006, 12 July 2006; http://gladiator.
+ncsa.uiuc.edu/PDFs/rssi06/presentations/00_Duncan_Buell.pdf
+[3] T. Van Court et al., Achieving High Performance
+with FPGA-Based Computing, Computer, vol. 40, no. 3, 
+pp. 50-57, Mar. 2007, doi:10.1109/MC.2007.79
+\end{verbatim}
+\end{scriptsize}
+
+\subsubsection{System Synthesis}
+Today, several solutions for system design are proposed and commercialized. The most common are
+those provided by Altera and Xilinx to promote their FPGA devices.
+\\
+The Xilinx System Generator for DSP [http://www.xilinx.com/tools/sysgen.htm] is a plug-in to 
+Simulink that enables designers to develop high-performance DSP systems for Xilinx FPGAs. 
+Designers can design and simulate a system using MATLAB and Simulink. The tool will then 
+automatically generate synthesizable Hardware Description Language (HDL) code mapped to Xilinx 
+pre-optimized algorithms. 
+However, this tool targets only DSP based algorithms, Xilinx FPGAs and cannot handle complete
+SoC. Thus, it is not really a system synthesis tool.
+\\
+In the opposite, SOPC Builder [CITATION] allows to describe a system, to synthesis it, 
+to programm it into a target FPGA and to upload a software application. 
+% FIXME(C2H from Altera, marche vite mais ressource monstrueuse)
+Nevertheless, SOPC Builder does not provide any facilities to synthesize coprocessors.
+Users have to provide the synthesizable description with the feasible bus interface.
+\\
+In addition, Xilinx System Generator and SOPC are closed world since each one imposes
+their own IPs which are not interchangeable.
+We can conclude that the existing commercial or free tools does not coverthe whole system 
+synthesis process in a full automatic way. Moreover, they are bound to a particular device family
+and to IPs library.
+
+\subsubsection{High Level Synthesis}
+High Level Synthesis translates a sequential algorithmic description and a constraints set 
+(area, power, frequency, ...) to a micro-architecture at Register Transfer Level (RTL).
+Several academic and commercial tools are today available. 
+Most common tools are SPARK [HLS1], GAUT [HLS2], UGH [HLS3] in the academic world 
+and catapultC [HLS4], PICO [HLS5] and Cynthesizer [HLS6] in commercial world.
+Despite their maturity, their usage is restrained by:
+\begin{itemize}
+\item They do not respect accurately the frequency constraint when they target an FPGA device.
+Their error is about 10 percent. This is annoying when the generated component is integrated
+in a SoC since it will slow down the hole system.
+\item These tools take into account only one or few constraints simultaneously while realistic
+designs are multi-constrained. 
+Moreover, low power consumption constraint is mandatory for embedded systems. 
+However, it is not yet well handled by common synthesis tools.
+\item The parallelism is extracted from initial algorithm. To get more parallelism or to reduce
+the amout of required memory, the user must re-write it while there is techniques as polyedric 
+transformations to increase the intrinsec parallelism.
+\item Despite they have the same input language (C/C++), they are sensitive to the style in
+which the algorithm is written. Consequently, engineering work is required to swap from 
+a tool to another.
+\item The HLS tools are not integrated into an architecture and system exploration tool.
+Thus, a designer who needs to accelerate a software part of the system, must adapt it manually 
+to the HLS input dialect and performs engineering work to exploit the synthesis result 
+at the system level.
+\end{itemize}
+Regarding these limitations, it is necessary to create a new tool generation reducing the gap 
+between the specification of an heterogenous system and its hardware implementation.
+
+\hspace{2cm}\begin{scriptsize}\begin{verbatim}
+[HLS1] SPARK universite de californie San Diego
+[HLS2] GAUT UBS/Lab-STIC
+[HLS3] UGH
+[HLS4] catapultC Mentor
+[HLS5] PICO synfora
+[HLS6] Cynthesizer Forte design system 
+\end{verbatim}
+\end{scriptsize}
+
+\subsubsection{Application Specific Instruction Processors}
+
+ASIP (Application-Specific Instruction-Set Processor) are programmable processors in 
+which both the instruction and the micro architecture have been tailored to a given
+ application domain (eg. video processing), or to a specific application. 
+This specialization usually offers a good compromise between performance (w.r.t a pure software
+implementation on an embeded CPU) and flexibility (w.r.t an application specific 
+hardware co-processor).
+In spite of their obvious advantages, using/designing ASIPs remains a difficult
+task, since it involves designing both a micro-architecture and a compiler for this
+architecture. Besides, to our knowledge, there is still no available open-source
+design flow\footnote{There are commercial tools such a } for ASIP design even if such a tool would
+be valuable in the context of a System Level design exploration tool.    
+
+In this context, ASIP design based on Instruction Set Extensions (ISEs) has 
+received a lot of interest [NIOSII,TENSILICA]%~\cite{NIOS2,ST70}, 
+as it makes micro architecture synthesis 
+more tractable \footnote{ISEs rely on a template micro-architecture in which 
+only a small fraction of the architecture has to be specialized}, and help ASIP
+designers to focus on compilers, for which there are still many open problems 
+[CODES04,FPGA08].
+This approach however has a strong weakness, since it also significantly reduces 
+opportunities for achieving good seedups (most speedup remain between 1.5x and 
+2.5x), since ISEs performance is generally tied down by I/O constraints as 
+they generally rely on the main CPU register file to access data.
+
+% (
+%automaticcaly extraction ISE candidates for application code \cite{CODES04}, 
+%performing efficient instruction selection and/or storage resource (register) 
+%allocation \cite{FPGA08}).  
+ 
+
+To cope with this issue, recent approaches~[DAC09,DAC08]%\cite{DAC09,DAC08} 
+advocate the use of 
+micro-architectural ISE models in which the coupling between the processor micro-architecture
+and the ISE component is thightened up so as to allow the ISE to overcome the register 
+I/O limitations, however these approaches tackle the problem for a compiler/simulation 
+point of view and not address the problem of generating synthesizable representations for 
+these models. 
+
+We therefore strongly believe that there is a need for an open-framework which
+would allow researchers and system designers to :
+\begin{itemize}
+\item Explore the various level of interactions between the original CPU micro-architecure
+and its extension (for example throught a Domain Specific Language targeted at micro-architecture
+specification and synthesis).
+\item Retarget the compiler instruction-selection (or prototype nex passes) passes so as
+to be able to take advantage of this ISEs.
+\item Provide  a complete System-level Integration for using ASIP as SoC building blocks 
+(integration with application specific blocks, MPSoc, etc.)
+\end{itemize}
+
+\hspace{2cm}
+\begin{scriptsize}\begin{verbatim} 
+
+[CODES08] Theo Kluter, Philip Brisk, Paolo Ienne, and Edoardo Charbon, Speculative DMA for
+Architecturally Visible Storage in Instruction Set Extensions
+
+[DAC09] Theo Kluter, Philip Brisk, Paolo Ienne, Edoardo Charbon, Way Stealing: Cache-assisted
+Automatic Instruction Set Extensions.
+
+[CODES04] Pan Yu, Tulika Mitra, Scalable Custom Instructions Identification for
+Instruction Set Extensible Processors.
+
+[FPGA08] Quang Dinh, Deming Chen, Martin D. F. Wong, Efficient ASIP Design for Configurable
+Processors with Fine-Grained Resource Sharing.
+
+[NIOSII] Nios II Custom Instruction User Guide
+
+\end{verbatim}
+
+\end{scriptsize}
+%, either 
+%because the target architecture is proprietary, or because the compiler 
+%technology is closed/commercial.
+
+
+
+
+% We propose to explore how to tighten the coupling of the extensions and 
+% the underlyoing template micro-architecture.
+% *  Thightne Even if such 
+% an approach offers less flexiblity and forbids very tight coupling 
+% between the extensions and the template micro-architecture, it makes the 
+% design of the micro-architecture more tractable and amenable to a fully 
+% automated flow.
+% \\
+% \\
+% In the context of the COACH project, we propose to add to the 
+% infra-structure a design flow targeted to automatic instruction set 
+% extension for the MIPS-based CPU, which will come as a complement or an 
+% alternative to the other proposed approaches (hardware accelerator, 
+% multi processors).
+% 
+
+\subsubsection{Automatic Parallelization}
+\begin{Large}\begin{verbatim}
+-- A COMPLETER LIP
+\end{verbatim}
+\end{Large}
+%CA%   Parallel machines are often difficult and painful to program
+%CA%   directly, and one would like the compiler to %do the job, that is to
+%CA%   turn automatically a sequential program into a parallel form. This
+%CA%   transformation is referred as {\em automatic parallelization}, and has
+%CA%   been widely addressed since the 70s. Automatic parallelization
+%CA%   relies on data dependences, which cannot be computed in general.%, as
+%CA%   %one cannot predict at compile time the variable values on a given
+%CA%   %execution point. 
+%CA%   This negative result led researchers to (i) find a
+%CA%   program model in which no approximation is needed (ie polyhedral
+%CA%   model), (ii) make conservative approximations (iii) remark that
+%CA%   variable values are known at runtime, and make the decisions during
+%CA%   program execution. The latter approach is obviously not suitable
+%CA%   there, as we target hardware generation. We will give there a short
+%CA%   history of the approaches that fall in the first category.
+%CA%
+%CA%%   In the real world, we deal with a limited amount of processors,
+%CA%%   and the communication between processors takes time, and is
+%CA%%   critical for performance. %Whenever we have synchronisation-free
+%CA%%   parallelism, like for embarrassingly parallel kernels, this is not an
+%CA%%   issue. But in case of pipelined parallelism, we need to reduce
+%CA%%   communications as much as possible. 
+%CA%%   So we also need to find parallelism toghether with a proper mapping
+%CA%%   of operations and data on physical processors.
+%CA%
+%CA%   As programs spend most of there time in loops, the community has
+%CA%   focused on loop transformations that reveal parallelism. 
+%CA%%unimodulaire
+%CA%   The first approaches worked on perfect loop nests, where the tree
+%CA%   formed by the nested loops is linear. In this program model, the
+%CA%   loops can be seen as a basis that drive the way the iteration
+%CA%   domain will be described. Hence, a first idea was to change this
+%CA%   basis such that one vector (one loop) at least is parallel. To ease
+%CA%   the code generation, the area of defined by the news vectors must
+%CA%   be a unit volume. %Otherwise, one would produce an homothetic
+%CA%%   expansion of the iteration domain, which will force to put modulos
+%CA%%   in the target code. 
+%CA%   For this reason, these transformations are called {\em unimodular
+%CA%   transformations}.
+%CA%%tiling
+%CA%   
+%CA%   The next approaches include {\em loop tiling}, a simple
+%CA%   partitioning of the iteration domain, whose initial purpose is to
+%CA%   execute every partition on a different processor. %In the same way,
+%CA%   The execution order is modified with a proper unimodular
+%CA%   transformation, then the tiles are obtained by cutting the
+%CA%   iteration domain with the hyperplanes directed by every vector of
+%CA%   the new (unimodular) basis, at regular intervals. When the tiling
+%CA%   hyperplanes are properly chosen, we can both improve data-locality
+%CA%   on every processor, and reduce the communication between two
+%CA%   different tiles (which will be mapped on processors). This last
+%CA%   property implying that one tend to find a degree of parallelism as
+%CA%   great as possible.
+%CA%
+%CA%%affine scheduling
+%CA%   The previous approaches were restricted to kernels with perfect
+%CA%   loop nests (linear loop tree), and unimodular transformations. The
+%CA%   last generation of approaches broke with these limitations. We now
+%CA%   choose a different basis for every assignment, without the
+%CA%   unimodularity restriction. A dual way to present the things is the
+%CA%   notion of {\em affine schedule}, introduced by Feautrier [part1],
+%CA%   that simply assigns an abstract execution date to every assignment
+%CA%   execution. As an assignment execution is exactly characterised by
+%CA%   the current value of the loops counters (iteration vector), the
+%CA%   affine schedule will be defined as an affine form of the iteration
+%CA%   vector (hence the 'affine'). The affine property allows to use
+%CA%   integer programming techniques to compute the schedule. With this
+%CA%   approach, additional techniques are required to allocate the
+%CA%   parallel operations and the data to processor in an efficient way
+%CA%   [griebl, feautrier].
+%CA%
+%CA%%modularity??
+%CA%%%    As loop nests are no longer perfect, we deal with (transformed)
+%CA%%%    iteration domains of different dimensions, which can possibly (and
+%CA%%%    certainly) overlap. At this point, a new code generation technique
+%CA%%%    was needed. The first attempt is due to Chamsky et al. [??], and
+%CA%%%    was improved by Quillere et al. [QRW]. The code is now implemented
+%CA%%%    in an efficient tool [cloog], that gave a new life to polyhedral
+%CA%%%    techniques.
+%CA%
+%CA%%pluto's tiling
+%CA%   The tiling techniques were extended to non-perfect loop nest with
+%CA%   {\em affine partitioning}. Affine partitioning is to affine
+%CA%   scheduling what (original) tiling was to unimodular
+%CA%   transformations. An affine partitioning assigns to every assignment
+%CA%   its coordinates in the basis defined by the normals to the tiling
+%CA%   hyperplanes. Recently, a way to compute efficient hyperplanes were
+%CA%   found [uday], with a good data locality, and communications
+%CA%   confined in a small neighborhood around every processor.
+%CA%
+%CA%\subsubsection{Source-level Memory Optimisation}
+%CA%  The HLS process allows to customise memory, which impacts on final
+%CA%  circuit size and power consumption. Though most HLS tools already
+%CA%  try to optimise memory usage, it is better to provide an independent
+%CA%  source-level pass, that could be reused for different tools and in
+%CA%  other contexts.
+%CA%
+%CA%  There exists many approaches to evaluate and reduce the memory
+%CA%  requirement of a program. The first approaches are concerned with
+%CA%  {\em memory size estimation}, which can be defined as the maximum
+%CA%  number of memory cells used at the same time [clauss,zhao]. These
+%CA%  approaches provide an estimation as a symbolic expression of program
+%CA%  parameters, which can be used further to guide loop optimisations.
+%CA%  However, no explicit way to reduce the memory size is given.  {\em
+%CA%  Intra-array reuse} approaches brake with this limitation, and
+%CA%  collapse the array cells which are not alive at the same time. The
+%CA%  collapse is done by means of a data layout transformation, specified
+%CA%  with a linear (modular) mapping.  The first approaches were
+%CA%  developed at IMEC [balasa,catthoor], and basically try to linearize
+%CA%  the arrays and fold them using a modulo operator. Then, Lefebvre et
+%CA%  al. propose a solution to fold independently the array dimensions
+%CA%  [lefebvre]. Finally, Darte et al. provide a general formalisation of
+%CA%  the problem, together with a solution that subsumes the previous
+%CA%  approaches [darte]. A first implementation was made with the tool
+%CA%  {\sc Bee}, but there are still many limitations.
+%CA%
+%CA%  \begin{itemize} 
+%CA%  \item The tool is restricted to regular programs, whereas more
+%CA%  general programs could be handled with a conservative array liveness
+%CA%  analysis.
+%CA%
+%CA%  \item Programs depending on parameters (inputs) are not handled,
+%CA%  which forbids to handle, for example, the body of tiled loops.
+%CA%
+%CA%  \item The new array layout can brake spatial locality, and then impact
+%CA%  performance and power consumption. One would like to get a mapping
+%CA%  that improve or, at least, preserve the spatial locality of the
+%CA%  program.
+%CA%
+%CA%  \item Finally, the final memory compaction strongly depends on the
+%CA%  program schedule, and is naturally hindered by the
+%CA%  parallelism. Consequently, there is a trade-off to find with
+%CA%  automatic parallelization. An ideal solution would be to reduce
+%CA%  memory usage, while preserving parallelism.  
+%CA%  \end{itemize}
+
+\subsubsection{Interfaces}
+\begin{Large}\begin{verbatim}
+-- A COMPLETER INSA Etat de l'art
+\end{verbatim}
+\end{Large}
+%
+%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
+\subsection{Objectives and innovation aspects}
+\hspace{2cm}\begin{scriptsize}\begin{verbatim}
+% 2.2.	OBJECTIFS ET CARACTERE AMBITIEUX/NOVATEUR DU PROJET 
+% (2 pages maximum)
+% Dï¿œcrire les objectifs scientifiques/techniques du projet.
+% Prï¿œsenter l'avancï¿œe scientifique attendue. Prï¿œciser l'originalitï¿œ et le caractï¿œre 
+% ambitieux du projet.
+% Dï¿œtailler les verrous scientifiques et techniques ï¿œ lever par la rï¿œalisation du projet.
+% Dï¿œcrire ï¿œventuellement le ou les produits finaux dï¿œveloppï¿œs ï¿œ l'issue du projet  
+% montrant le caractï¿œre innovant du projet.
+% Prï¿œsenter les rï¿œsultats escomptï¿œs en proposant si possible des critï¿œres de rï¿œussite 
+% et d'ï¿œvaluation adaptï¿œs au type de projet, permettant d'ï¿œvaluer les rï¿œsultats en 
+% fin de projet.
+% Le cas ï¿œchï¿œant (programmes exigeant la pluridisciplinaritï¿œ), dï¿œmontrer l'articulation 
+% entre les disciplines scientifiques.
+\end{verbatim}
+\end{scriptsize}
+
+% les objectifs scientifiques/techniques du projet.
+The objectives of COACH project are to develop a complete framework to 
+HPC (accelerating solutions for existing software applications)
+and embedded applications (implementing an application on a low power standalone device).
+The design steps are presented figure 1.
+\begin{figure}[hbtp]\leavevmode\center
+  \includegraphics[width=.8\linewidth]{flow}
+  \caption{\label{coach-flow} COACH flow.}
+\end{figure}
+\begin{description}
+\item[HPC setup] Here the user splits the application into 2 parts: the host application
+which remains on PC and the SoC application which migrates on SoC. 
+The framework provides a simulation model allowing to evaluate the partitioning.
+\item[SoC design] In this phase, 
+The user can obtain simulators at different abstraction levels of the SoC by giving to COACH framework
+a SoC description.  
+This description consists of a process network corresponding to the SoC application, 
+an OS, an instance of a generic hardware platform
+and a mapping of processes on the platform components. The supported mapping are 
+software (the process runs on a SoC processor),
+XXXpeci (the process runs on a SoC processor enhanced with dedicated instructions),
+and hardware (the process runs into a coprocessor generated by HLS and plugged on the SoC bus).
+\item[Application compilation] Once SoC description is validated, COACH generates automatically
+an FPGA bitstream containing the hardware platform with SoC application software and 
+an executable containing the host application. The user can launch the application by
+loading the bitstream on FPGA and running the executable on PC.
+\end{description}
+ 
+% l'avancee scientifique attendue. Preciser l'originalite et le caractere 
+% ambitieux du projet. 
+The main scientific contribution of the project is to unify various synthesis techniques
+(same input and output formats) allowing the user to swap without engineering effort
+from one to an other and even to chain them, for example, to run polyedric transformation 
+before synthesis.
+Another advantage of this framework is to provide different abstraction levels from
+a single description.
+Finally, this description is device family independent and its hardware implementation
+is automatically generated.
+
+% Detailler les verrous scientifiques et techniques a lever par la realisation du projet.
+System design is a very complicated task and in this project we try to simplify it
+as much as possible. For this purpose we have to deal with the following scientific
+and technological barriers.
+\begin{itemize}
+\item The main problem in HPC is the communication between the PC and the SoC.
+This problem has 2 aspects. The first one is the efficiency. The second is to 
+eliminate enginnering effort to implement it at different abstract levels.
+\item COACH design flow has a top-down approach. In the such case,
+the required performance of a coprocessor (run frequency, maximum cycles for
+a given computation, power consumption, etc) are imposed by the other system
+components. The challenge is to allow user to control accurately the synthesis
+process. For instance, the run frequency must not be a result of the RTL synthesis
+but a strict synthesis constraint.
+\item HLS tools are sensitive to the style in which the algorithm is written.
+In addition, they are are not integrated into an architecture and system 
+exploration tool.
+Consequently, engineering work is required to swap from a tool to another,
+to integrate the resulting simulation model to an architectural exploration tool 
+and to synthesize the generated RTL description.
+%CA Additionnal preprocessing, source-level transformations, are thus
+%CA required to improve the process.
+%CA Particularly, this includes parallelism exposure and efficient memory mapping.
+\item Most HLS tools translate a sequential algorithm into a coprocessor
+containing a single data-path and finite state machine (FSM). In this way,
+only the fine grained parallelism is exploited (ILP parallelism).
+The challenge is to identify the coarse grained parallelism and to generate,
+from a sequential algorithm, coprocessor containing multiple communicating
+tasks (data-paths and FSMs).
+\end{itemize}
+
+%Presenter les resultats escomptes en proposant si possible des criteres de reussite 
+%et d'evaluation adaptes au type de projet, permettant d'evaluer les resultats en 
+%fin de projet.
+The main result is the framework. It is composed concretely of: 
+2 HPC communication shemes with their implementation, 
+5 HLS tools (control dominated HLS, data dominated HLS, Coarse grained HLS, 
+Memory optimisation HLS and ASIP),
+3 systemC based virtual prototyping environment extended with synthesizable
+RTL IP cores (generic, ALTERA/NIOS/AVALON, XILINX/MICROBLAZE/OPB),
+one design space exploration tool,
+one operating system (OS).
+\\
+The framework fonctionality will be demonstrated with XXX-EXAMPLE1, XXX-EXAMPLE2
+and XXX-EXAMPLE3 on 4 archictures (generic/XILINX, generic/ALTERA,
+proprietary/XILINX, proprietary/ALTERA).
+
+%% \section{}
+%% %3.	PROGRAMME SCIENTIFIQUE ET TECHNIQUE, ORGANISATION DU PROJET
+%% \subsection{}
+%% %3.1.	PROGRAMME SCIENTIFIQUE ET STRUCTURATION DU PROJET 
+%% %(2 pages maximum)
+%% %Prï¿œsentez le programme scientifique et justifiez la dï¿œcomposition en tï¿œches du 
+%% %programme de travail en cohï¿œrence avec les objectifs poursuivis. 
+%% %Utilisez un diagramme pour prï¿œsenter les liens entre les diffï¿œrentes tï¿œches 
+%% %(organigramme technique)
+%% %Les tï¿œches reprï¿œsentent les grandes phases du projet. Elles sont en nombre limitï¿œ.
+%% %N'oubliez pas les activitï¿œs et actions correspondant ï¿œ la dissï¿œmination et ï¿œ la 
+%% %valorisation.
+%% 
+%% %METTRE UNE FIGURE ICI DECRIVANT LES TACHES ET LEURS INTERACTION (AVEC LE FLOT  
+%% %EN FILIGRANE ? )
+%% \subsection{}
+%% %3.2.	MANAGEMENT DU PROJET
+%% %(2 pages maximum)
+%% %Prï¿œciser les aspects organisationnels du projet et les modalitï¿œs de coordination 
+%% %(si possible individualisation d'une tï¿œche coordination : cf. tï¿œche 0 du document 
+%% %de soumission A).
+%% \subsection{}
+%% %3.3.	DESCRIPTION DES TRAVAUX PAR TACHE
+%% %(idï¿œalement 1 ou 2 pages par tï¿œche)
+%% %Pour chaque tï¿œche, dï¿œcrire : 
+%% %-	les objectifs  de la tï¿œche et ï¿œventuels indicateurs de succï¿œs,
+%% %-	le responsable de la tï¿œche et les partenaires impliquï¿œs (possibilitï¿œ de 
+%% %l'indiquer sous forme graphique),
+%% %-	le programme dï¿œtaillï¿œ des travaux par tï¿œche,
+%% %-	les livrables de la tï¿œche,
+%% %-	les contributions des partenaires (le " qui fait quoi "),
+%% %-	la description des mï¿œthodes et des choix techniques et de la maniï¿œre dont 
+%% %les solutions seront apportï¿œes,
+%% %-	les risques de la tï¿œche et les solutions de repli envisagï¿œes.
+
+
+
+
+
+
Index: /anr/obsolete/coach_irisa.bib
===================================================================
--- /anr/obsolete/coach_irisa.bib	(revision 12)
+++ /anr/obsolete/coach_irisa.bib	(revision 12)
@@ -0,0 +1,33 @@
+@InProceedings{KluterCodes08,
+  author = 	 {{Theo Kluter and  Philip Brisk and  Paolo Ienne and  and Edoardo Charbon}},
+  title = 	 {{Speculative DMA for Architecturally Visible Storage in Instruction Set Extensions}},
+  booktitle = {ISSS/CODES},
+  year = 	 {2008},
+}
+
+@InProceedings{KluterDAC09,
+  author = 	 {{Theo Kluter and  Philip Brisk and  Paolo Ienne and  and Edoardo Charbon}},
+  title = 	 {{Way Stealing : Cache-assisted Automatic Instruction Set Extensions}},
+  booktitle = {Design Automation Conference (DAC)},
+  year = 	 {2009},
+}
+
+@InProceedings{YuCodes04,
+  author = 	 {{Pan Yu and Tulika Mitra}},
+  title = 	 {{Scalable Custom Instructions Identification for Instruction Set Extensible Processors}},
+  booktitle = {ISSS/CODES},
+  year = 	 {2004},
+}
+
+@InProceedings{Dinh08,
+  author = 	 {{Quang Dinh and Deming Chen and Martin D.~F.~Wong}},
+  title = 	 {{Efficient ASIP Design for Configurable Processors with Fine-Grained Resource Sharing}},
+  booktitle = {ACM Internatibnal Conference Field Programmable Gate Arrays (FPGA)},
+  year = 	 {2008},
+}
+
+@Misc{NIOS2UG,
+  title = 	 {{Nios II Custom Instruction User Guide, Altera Corp.}},
+  year = 	 {2008},
+  
+}
Index: /anr/obsolete/wp-IRISA.tex
===================================================================
--- /anr/obsolete/wp-IRISA.tex	(revision 12)
+++ /anr/obsolete/wp-IRISA.tex	(revision 12)
@@ -0,0 +1,79 @@
+\documentclass[11pt,a4paper]{article}
+
+\usepackage[french]{babel}
+\usepackage[utf8x]{inputenc}
+\usepackage{times}
+\usepackage[T1]{fontenc}
+\usepackage{aeguill}
+\usepackage{verbatim}
+\usepackage{algorithm,algorithmic}
+\usepackage{xmpmulti}
+\usepackage{graphicx}
+\usepackage{color}
+
+\definecolor{gris25}{gray}{0.75}
+\definecolor{gris75}{gray}{0.30}
+
+
+\begin{document}
+
+\section {IRISA-WP}
+
+\subsection{Work Package 1 : Un compilateur reciblable pour MIPS Ã©tendu}
+
+DÃ©livrable : software
+ 
+ImplÃ©mentation d'un back-end de compilation ciblant une version Ã©tendue du processeurs MIPS, et
+opÃ©rant Ã  partir de la reprÃ©sentation intermÃ©diaire commune dÃ©finie en \ref{?} et issue de GCC. 
+Ce \emph{back-end} intÃšgrera en particulier des passes d'extraction de motifs de calculs
+(sous-graphes), ainsi une passe de sÃ©lection d'intructions basÃ©e sur des techniques de couvertures 
+de graphes, permettant d'exploiter au mieux les motifs d'instruction ``spÃ©cialisÃ©s'' spÃ©cifiÃ©s 
+par l'utilisateur et/ou extraits Ã  partir de l'application.
+\textcolor{gris75}{Ici, il faut voir sir la RI prposÃ©e ne permettrait pas de regÃ©nÃ©rer un code 
+C dans lequel l'utilisation d'instruction spÃ©cialisÃ©e se fait au travers de directives de type
+\texttt{asm\{ \ldots \}} . une telle approche offrirait une flexibilitÃ© accrue, sans 
+impacter la qualitÃ©/performance des rÃ©sultats obtenus.}
+
+\subsection{Work Package 2 : DÃ©finition d'un modÃšle simplifiÃ© de micro-architecture MIPS extensible}
+
+DÃ©livrable : software
+
+DÃ©finition d'un modÃšle extensible de micro-architecture basÃ©e sur un processeur de type MIPS
+pipelinÃ© Ã  5 Ã©tage (incluant cache de donnÃ©es et instruction), offrant Ã  l'utilisateur 
+la possibilitÃ© de dÃ©finir ses propres extensions architecturales, au travers d'un Domain
+Specific Language (on exploitera les technologies d'IDM XText-EMF). 
+
+On mettra Ã©galement en {\oe}uvre un outil de gÃ©nÃ©ratiion de description matÃ©rielle synthÃ©tisable
+(VHDL) de la micro-architecture Ã  partir de ce modÃšle, en utilisant des techonologies d'IngÃ©nierie 
+dirigÃ©e par les modÃšles (EMF-XPAND)
+ 
+Deux version de cet outil sont envisagÃ©es, dans la premiÃšre (qui est l'object de ce WP), on
+restreindra les possibilitÃ©s de communication entre le processeur et ses extension Ã  des communictaion passant par la file de
+regsitre du processeur (en permettant Ã©ventuellement un plus grand nombre d'accÃšs en
+lecture.Ã©criture par cycle).
+
+\subsection{Work Package 3 : DÃ©finition d'un modÃšle complexe de micro-architecture MIPS extensible}
+
+DÃ©livrable : rapport/software ?
+
+Dans la seconde version (plus orientÃ©e exploratoire) on souhaite pouvoir lever la limitaion portant
+sur les communications  et permettre un couplage plus fin entre les extensions et le coeur de la
+micro-rachgitcture, par exemple, en proposant un accÃšs direct au cache de donnÃ©es et/ou en donnant
+la possibiliteÃ© aux extensions de rÃ©utiliser les opÃ©rateurs (par exemple le multiplieur
+$32 \times 32$ \, bits) mis en oeuvre dans le chemin de donnÃ©e natif du processeur.
+
+Ici on pourrait Ã©galement envisager des connections directes avec d'autres composants au travers de
+structures similaire aux \emph{FSL} disponibles sur les processeurs softcore Microblaze de la
+sociÃ©tÃ© Xilinx.
+
+\subsection{Work Package 4 : DÃ©finition d'un modÃšle complexe de micro-architecture MIPS extensible}
+
+DÃ©livrable : rapport 
+
+Le dernier \emph{package} a Ã©galement un caractÃšre exploratoire, et portera sur l'intÃ©gration de ce
+type d'extension architecturales au sein d'un compilateur. En particulier, il s'agira d'Ã©tudier
+comment il est possible d'intÃ©grer ces instructions complexes dans la passe de selection de code,
+tout en s'assurant que leur contraintes d'utilisation soient respectÃ©es.
+
+
+\end{document}
Index: /anr/section-3.1.tex
===================================================================
--- /anr/section-3.1.tex	(revision 12)
+++ /anr/section-3.1.tex	(revision 12)
@@ -0,0 +1,219 @@
+Our project covers several critical domains in system design in order
+to achieve high performance computing. Starting from a high level description we aim 
+at generating automatically both hardware and software components of the system.
+
+\subsubsection{High Performance Computing}
+Accelerating high-performance computing (HPC) applications with field-programmable
+gate arrays (FPGAs) can potentially improve performance. 
+However, using FPGAs presents significant challenges~\cite{hpc06a}.
+First, the operating frequency of an FPGA is low compared to a high-end microprocessor.
+Second, based on Amdahl law,  HPC/FPGA application performance is unusually sensitive 
+to the implementation quality~\cite{hpc06b}.
+Finally, High-performance computing programmers are a highly sophisticated but scarce 
+resource. Such programmers are expected to readily use new technology but lack the time 
+to learn a completely new skill such as logic design~\cite{hpc07a} . 
+\\
+HPC/FPGA hardware is only now emerging and in early commercial stages, 
+but these techniques have not yet caught up. 
+Thus, much effort is required to develop design tools that translate high level
+language programs to FPGA configurations.
+
+\subsubsection{System Synthesis}
+Today, several solutions for system design are proposed and commercialized.
+The most common are those provided by Altera and Xilinx to promote their
+FPGA devices.
+\\
+The Xilinx System Generator for DSP~\cite{system-generateur-for-dsp} is a
+plug-in to Simulink that enables designers to develop high-performance DSP
+systems for Xilinx FPGAs.
+Designers can design and simulate a system using MATLAB and Simulink. The
+tool will then automatically generate synthesizable Hardware Description
+Language (HDL) code mapped to Xilinx pre-optimized algorithms.
+However, this tool targets only DSP based algorithms, Xilinx FPGAs and
+cannot handle complete SoC. Thus, it is not really a system synthesis tool.
+\\
+In the opposite, SOPC Builder~\cite{spoc-builder} allows to describe a
+system, to synthesis it, to programm it into a target FPGA and to upload a
+software application.
+% FIXME(C2H from Altera, marche vite mais ressource monstrueuse)
+Nevertheless, SOPC Builder does not provide any facilities to synthesize
+coprocessors. System Designer must provide the synthesizable description
+with the feasible bus interface.
+\\
+In addition, Xilinx System Generator and SOPC Builder are closed world
+since each one imposes their own IPs which are not interchangeable.
+We can conclude that the existing commercial or free tools does not
+coverthe whole system synthesis process in a full automatic way. Moreover,
+they are bound to a particular device family and to IPs library.
+
+\subsubsection{High Level Synthesis}
+High Level Synthesis translates a sequential algorithmic description and a
+constraints set (area, power, frequency, ...) to a micro-architecture at
+Register Transfer Level (RTL).
+Several academic and commercial tools are today available. Most common
+tools are SPARK~\cite{spark04}, GAUT~\cite{gaut08}, UGH~\cite{ugh08} in the
+academic world and CATAPULTC~\cite{catapult-c}, PICO~\cite{pico} and
+CYNTHETIZER~\cite{cynthetizer} in commercial world.  Despite their
+maturity, their usage is restrained by:
+\begin{itemize}
+\item They do not respect accurately the frequency constraint when they target an FPGA device.
+Their error is about 10 percent. This is annoying when the generated component is integrated
+in a SoC since it will slow down the hole system.
+\item These tools take into account only one or few constraints simultaneously while realistic
+designs are multi-constrained. 
+Moreover, low power consumption constraint is mandatory for embedded systems. 
+However, it is not yet well handled by common synthesis tools.
+\item The parallelism is extracted from initial algorithm. To get more parallelism or to reduce
+the amout of required memory, the user must re-write it while there is techniques as polyedric 
+transformations to increase the intrinsec parallelism.
+\item Despite they have the same input language (C/C++), they are sensitive to the style in
+which the algorithm is written. Consequently, engineering work is required to swap from 
+a tool to another.
+\item The HLS tools are not integrated into an architecture and system exploration tool.
+Thus, a designer who needs to accelerate a software part of the system, must adapt it manually 
+to the HLS input dialect and performs engineering work to exploit the synthesis result 
+at the system level.
+\end{itemize}
+Regarding these limitations, it is necessary to create a new tool generation reducing the gap 
+between the specification of an heterogenous system and its hardware implementation.
+
+\subsubsection{Application Specific Instruction Processors}
+
+ASIP (Application-Specific Instruction-Set Processor) are programmable
+processors in which both the instruction and the micro architecture have
+been tailored to a given application domain (eg. video processing), or to a
+specific application.  This specialization usually offers a good compromise
+between performance (w.r.t a pure software implementation on an embeded
+CPU) and flexibility (w.r.t an application specific hardware co-processor).
+In spite of their obvious advantages, using/designing ASIPs remains a
+difficult task, since it involves designing both a micro-architecture and a
+compiler for this architecture. Besides, to our knowledge, there is still
+no available open-source design flow\footnote{There are commercial tools
+such a } for ASIP design even if such a tool would be valuable in the
+context of a System Level design exploration tool.
+\par
+In this context, ASIP design based on Instruction Set Extensions (ISEs) has 
+received a lot of interest~\cite{NIOS2,ST70}, as it makes micro architecture synthesis 
+more tractable \footnote{ISEs rely on a template micro-architecture in which 
+only a small fraction of the architecture has to be specialized}, and help ASIP
+designers to focus on compilers, for which there are still many open
+problems\cite{CODES04,FPGA08}.
+This approach however has a strong weakness, since it also significantly reduces 
+opportunities for achieving good seedups (most speedup remain between 1.5x and 
+2.5x), since ISEs performance is generally tied down by I/O constraints as 
+they generally rely on the main CPU register file to access data.
+
+% (
+%automaticcaly extraction ISE candidates for application code \cite{CODES04}, 
+%performing efficient instruction selection and/or storage resource (register) 
+%allocation \cite{FPGA08}).  
+To cope with this issue, recent approaches~\cite{DAC09,DAC08} advocate the use of 
+micro-architectural ISE models in which the coupling between the processor micro-architecture
+and the ISE component is thightened up so as to allow the ISE to overcome the register 
+I/O limitations, however these approaches tackle the problem for a compiler/simulation 
+point of view and not address the problem of generating synthesizable representations for 
+these models. 
+
+We therefore strongly believe that there is a need for an open-framework which
+would allow researchers and system designers to :
+\begin{itemize}
+\item Explore the various level of interactions between the original CPU micro-architecure
+and its extension (for example throught a Domain Specific Language targeted at micro-architecture
+specification and synthesis).
+\item Retarget the compiler instruction-selection (or prototype nex passes) passes so as
+to be able to take advantage of this ISEs.
+\item Provide  a complete System-level Integration for using ASIP as SoC building blocks 
+(integration with application specific blocks, MPSoc, etc.)
+\end{itemize}
+
+\subsubsection{Automatic Parallelization}
+% FIXME:LIP FIXME:PF FIXME:CA
+% Paul je ne suis pas sur que ce soit vraiment un etat de l'art
+% Christophe, ce que tu m'avais envoye se trouve dans obsolete/body.tex
+\mustbecompleted{
+Hardware is inherently parallel. On the other hand, high level languages, 
+like C or Fortran, are abstractions of the processors of the 1970s, and
+hence are sequential. One of the aims of an HLS tool is therefore to
+extract hidden parallelism from the source program, and to infer enough
+hardaware operators for its efficient exploitation.
+\\
+Present day HLS tools search for parallelism in linear pieces of code
+acting only on scalars -- the so-called basic blocs. On the other hand,
+it is well known that most programs, especially in the fields of signal
+processing and image processing, spend most of their time executing loops
+acting on arrays. Efficient use of the large amount of hardware available
+in the next generation of FPGA chips necessitates parallelism far beyond
+what can be extracted from basic blocs only.
+\\
+The Compsys team of LIP has built an automatic parallelizer, Syntol, which
+handle restricted C programs -- the well known polyhedral model --, 
+computes dependences and build a symbolic schedule. The schedule is
+a specification for a parallel program. The parallelism itself can be 
+expressed in several ways: as a system of threads, or as data-parallel
+operations, or as a pipeline. In the context of the COACH project, one
+of the task will be to decide which form of parallelism is best suited
+to hardware, and how to convey the results of Syntol to the actual
+synthesis tools. One of the advantages of this approach is that the
+resulting degree of parallelism can be easilly controlled, e.g. by 
+adjusting the number of threads, as a mean of exploring the 
+area / performance tradeoff of the resulting design.
+\\
+Another point is that potentially parallel programs necessarily involve
+arrays: two operations which write to the same location must be executed
+in sequence. In synthesis, arrays translate to memory. However, in FPGAs,
+the amount of on-chip memory is limited, and access to an external memory
+has a high time penalty. Hence the importance of reducing the size of
+temporary arrays to the minimum necessary to support the requested degree
+of parallelism. Compsys has developped a stand-alone tool, Bee, based
+on research by A. Darte, F. Baray and C. Alias, which can be extended
+into a memory optimizer for COACH.
+}
+
+\subsubsection{Interfaces}
+\newcommand{\ip}{\sc ip}
+\newcommand{\dma}{\sc dma}
+\newcommand{\soc}{\sc SoC}
+\newcommand{\mwmr}{\sc mwmr}
+The hardware/software interface has been a difficult task since the advent
+of complex systems on chip. After the first Co-design
+environments~\cite{Coware,Polis,Ptolemy}, the Hardware Abstraction Layer
+has been defined so that software applications can be developed without low
+level hardware implementation details.  In~\cite{jerraya}, Yoo and Jerraya
+propose an {\sc api} with extension ability instead of a unique hardware
+abstraction layer.  System level communication frameworks have been
+introduced~\cite{JerrayaPetrot,mwmr}.
+\par
+A good abstraction of a hardware/software interface has been proposed
+in~\cite{Jantsch}: it is composed of a software driver, a {\dma} and and a
+bus interface circuit. Automatic wrapping between bus protocols has
+generated a lot of papers~\cite{Avnit,smith,Narayan, Alberto}. These works
+do not use a {\dma}. In COACH, the hardware/software interface is done at a
+higher level and uses burst communication in the bus interface circuit to
+improve the communication performances.
+\par
+There are two important projects related to efficient interface of
+data-flow {\ip}s : the work of Park and Diniz~\cite{ Park01} and the the
+Lip6 work on {\mwmr}~\cite{mwmr}.  Park and Diniz~\cite{ Park01} proposed
+of a generic interface that can be parameterized to connect different
+data-flow {\ip}s. This approach does not request the communications to be
+statically known and proposes a runtime resolution to solve conflicting
+access to the bus. To our knowledge this approach has not been implemented
+further since 2003.
+\par
+{\mwmr}~\cite{mwmr} stands for both a computation model (multi-write,
+multi-read {\sc fifo}) inherited from the Khan Process Networks and a bus
+interface circuit protocol.  As for the work of Park and Diniz, {\mwmr}
+does not make the assumption of a static communication flow.  This implies
+simple software driver to write, but introduces additional complexity due
+to the mutual exclusion locks necessary to protect the shared memory.
+\par
+we propose, in COACH, to use recent work on hardware/software
+interface~\cite{FR-vlsi}  that  uses a {\em clever} {\dma} responsible for
+managing data streams. A assumption is that the behavior of the {\ip}s can
+be statically described. A similar choice has been made in the Faust
+{\soc}~\cite{FAUST} which includes the {\em smart memory engine} component.
+Jantsch and O'Nils already noticed in ~\cite{Jantsch} the huge complexity
+of writing this hardware/software interface, in COACH,  automatic
+generation of the interface will be achieved, this is one goal of the CITI
+contribution to COACH.
+
Index: /anr/section-3.2.tex
===================================================================
--- /anr/section-3.2.tex	(revision 12)
+++ /anr/section-3.2.tex	(revision 12)
@@ -0,0 +1,86 @@
+% les objectifs scientifiques/techniques du projet.
+The objectives of COACH project are to develop a complete framework to 
+HPC (accelerating solutions for existing software applications)
+and embedded applications (implementing an application on a low power standalone device).
+The design steps are presented figure 1.
+\begin{figure}[hbtp]\leavevmode\center
+  \includegraphics[width=.8\linewidth]{flow}
+  \caption{\label{coach-flow} COACH flow.}
+\end{figure}
+\begin{description}
+\item[HPC setup] Here the user splits the application into 2 parts: the host application
+which remains on PC and the SoC application which migrates on SoC. 
+The framework provides a simulation model allowing to evaluate the partitioning.
+\item[SoC design] In this phase, 
+The user can obtain simulators at different abstraction levels of the SoC by giving to COACH framework
+a SoC description.  
+This description consists of a process network corresponding to the SoC application, 
+an OS, an instance of a generic hardware platform
+and a mapping of processes on the platform components. The supported mapping are 
+software (the process runs on a SoC processor),
+XXXpeci (the process runs on a SoC processor enhanced with dedicated instructions),
+and hardware (the process runs into a coprocessor generated by HLS and plugged on the SoC bus).
+\item[Application compilation] Once SoC description is validated, COACH generates automatically
+an FPGA bitstream containing the hardware platform with SoC application software and 
+an executable containing the host application. The user can launch the application by
+loading the bitstream on FPGA and running the executable on PC.
+\end{description}
+ 
+% l'avancee scientifique attendue. Preciser l'originalite et le caractere 
+% ambitieux du projet. 
+The main scientific contribution of the project is to unify various synthesis techniques
+(same input and output formats) allowing the user to swap without engineering effort
+from one to an other and even to chain them, for example, to run polyedric transformation 
+before synthesis.
+Another advantage of this framework is to provide different abstraction levels from
+a single description.
+Finally, this description is device family independent and its hardware implementation
+is automatically generated.
+
+% Detailler les verrous scientifiques et techniques a lever par la realisation du projet.
+System design is a very complicated task and in this project we try to simplify it
+as much as possible. For this purpose we have to deal with the following scientific
+and technological barriers.
+\begin{itemize}
+\item The main problem in HPC is the communication between the PC and the SoC.
+This problem has 2 aspects. The first one is the efficiency. The second is to 
+eliminate enginnering effort to implement it at different abstract levels.
+\item COACH design flow has a top-down approach. In the such case,
+the required performance of a coprocessor (run frequency, maximum cycles for
+a given computation, power consumption, etc) are imposed by the other system
+components. The challenge is to allow user to control accurately the synthesis
+process. For instance, the run frequency must not be a result of the RTL synthesis
+but a strict synthesis constraint.
+\item HLS tools are sensitive to the style in which the algorithm is written.
+In addition, they are are not integrated into an architecture and system 
+exploration tool.
+Consequently, engineering work is required to swap from a tool to another,
+to integrate the resulting simulation model to an architectural exploration tool 
+and to synthesize the generated RTL description.
+%CA Additionnal preprocessing, source-level transformations, are thus
+%CA required to improve the process.
+%CA Particularly, this includes parallelism exposure and efficient memory mapping.
+\item Most HLS tools translate a sequential algorithm into a coprocessor
+containing a single data-path and finite state machine (FSM). In this way,
+only the fine grained parallelism is exploited (ILP parallelism).
+The challenge is to identify the coarse grained parallelism and to generate,
+from a sequential algorithm, coprocessor containing multiple communicating
+tasks (data-paths and FSMs).
+\end{itemize}
+
+%Presenter les resultats escomptes en proposant si possible des criteres de reussite 
+%et d'evaluation adaptes au type de projet, permettant d'evaluer les resultats en 
+%fin de projet.
+The main result is the framework. It is composed concretely of: 
+2 HPC communication shemes with their implementation, 
+5 HLS tools (control dominated HLS, data dominated HLS, Coarse grained HLS, 
+Memory optimisation HLS and ASIP),
+3 systemC based virtual prototyping environment extended with synthesizable
+RTL IP cores (generic, ALTERA/NIOS/AVALON, XILINX/MICROBLAZE/OPB),
+one design space exploration tool,
+one operating system (OS).
+\\
+The framework fonctionality will be demonstrated with XXX-EXAMPLE1, XXX-EXAMPLE2
+and XXX-EXAMPLE3 on 4 archictures (generic/XILINX, generic/ALTERA,
+proprietary/XILINX, proprietary/ALTERA).
+
Index: /anr/section-4.1.tex
===================================================================
--- /anr/section-4.1.tex	(revision 12)
+++ /anr/section-4.1.tex	(revision 12)
@@ -0,0 +1,105 @@
+\begin{figure}\leavevmode\center
+\includegraphics[width=.8\linewidth]{architecture-csg}
+\caption{\label{archi-csg} software architecture for embedded system generation}
+%\end{figure}\begin{figure}\leavevmode\center
+\mbox{}\vspace*{1ex}\\
+\includegraphics[width=.8\linewidth]{architecture-hls}
+\caption{\label{archi-hls} software architecture of HLS}
+%\end{figure}\begin{figure}\leavevmode\center
+\mbox{}\vspace*{1ex}\\
+\includegraphics[width=.8\linewidth]{architecture-hpc}
+\caption{\label{archi-hpc} software architecture of HPC}
+\end{figure}
+%
+The figures~\ref{archi-csg}, \ref{archi-hls} and \ref{archi-hpc}
+summarize the software architecture of COACH framework we plan to develop.
+In figures, the dotted boxes are the softwares or formats that COACH
+has to provide or define.
+\vspace*{.75ex}\par
+For the system genration presented figure~\ref{archi-csg}, the conductor is
+the program \verb!CSG! (COACH System Generator). Its inputs are a process
+network and miscellaneaous generation parameters.
+The main parameters are the template of the target hardware architecture
+with its instanciation parameters, the hardware/software mapping of the
+tasks and the FPGA device.
+From these inputs \verb!CSG! can generate the system (software \& hardware) as
+a SystemC simulator to prototype and explore quickly the system design
+space and/or as a bitstream directly downloadable on the FPGA device.
+For processing, \verb+CSG+ requires 1) a hardware template found into the
+architecture library, 2) a micro-kernel, it chooses among
+two in the micro kernel library, 3) the system hardware components that
+are taken from the SystemC model library for the simulator and from the
+VHDL component library for the FPGA bitstream.  
+For generating the coprocessor of a task mapped as harware, \verb+CSG+
+controls the HLS tools described below.
+\\
+To proove CSG that COACH is open and CSG is really configurable, COACH will
+basically support 3 architecture template (the COACH template based on a
+MIPS processors and a VCI token ring, the Altera template based on the NIOS
+and AVALON bus, the Xilinx template based on the MICROBLAZE and OPB bus)
+and 2 operating systems (DNA/OS and MUTEK). Furthermore, thus is enforced
+by the \mustbecompleted{FIXME:zied} contribution that consists in
+implementing an other hardware target.
+\\
+Finally, it is important to notice that this work is a strong
+enhancement of the SocLib software.
+\vspace*{.75ex}\par
+The software architecture for HLS is presented figure~\ref{archi-hls}.
+The input is a task of the process network. The HLS tools do not work
+directly on the C++ task description but on an internal format called
+\xcoach generated by a the GNU C compiler (GCC) tainted by a COACH
+driver. This allows on the one hand to insure that all the tools will
+accept the same C++ description and on the other hand to make possible
+to chain them. The front-end tools read a \xcoach description and writes
+a new \xcoach description that exibits possible parallelism or implement
+specific instruction for ASIP. The back-end tools read a \xcoach
+description and generates a \xcoach+ description that is a \xcoach
+description anotated with hardware information to let work the VHDL systemC
+drivers. Furthermore, the back-end tools uses a macro-cell library.
+\vspace*{.75ex}\par
+The software architecture for HPC is presented figure~\ref{archi-hpc}.
+\mustbecompleted{FIXME Miss HPC description\\\ldots\\\ldots\\\ldots\\\ldots.}
+\vspace*{.75ex}\par
+The project is splitted into 8 tasks numbered from 0 to 7.
+The first task (task 0) is the project management, the last one (task 7) is
+the dissemination the other task are listed below:
+\begin{enumerate}
+\item\textbf{\backbone:} This task groups the critical issues of the
+	project. They consist of the definition of COACH inputs, the \xcoach
+	format that is mandatory to develop the HLS tools and the HLS drivers
+	that are mandatory for testing the HLS tools. 
+\item\textbf{system generation:} This task groups \verb+CSG+s and the components
+	required to generate the system simulator and bitstream except the HLS
+	tools that belong to the task 3 and 4. These components are the
+	operating systems, the VHDL description and SystemC models of the
+	target hardware achitectures.
+\item\textbf{HLS front-end:} This task groups the 4 HLS front-head. Those
+	are a tool that exhibits fine grain parallelism using polyedric
+	transformation, a tool that exhibits coarse grain parallelism,
+	a tool that minimizes the memory usage and a tool that implement ASIP.
+\item\textbf{HLS back-end:} This task groups two HLS back-end tool, one
+	for treating the data oriented description, the second for treating the
+	control dominated description. This task contains also a the
+	development of a frequency adaptator that will allow the coprocessor
+	to respect the processor \& bus frequency.
+\item\textbf{Communication software PC/FPGA-SoC:} This task groups all what is mandatorythe critical issues of the
+\item\textbf{Demonstrator:} This task groups the demostrators of the COACH project.
+\end{enumerate}
+This task division offers the avantage that tasks except for the "\backbone" and
+"demonstrator" task are almost independent at the development level as shown
+figure~\ref{dependence-dev}. The dependence at the validation level is
+presented figure~\ref{dependence-test}. It is more critical but the
+redundance in the tasks "HLS front-end", "HLS back-end" and "demonstrators"
+reduces this inter-dependence.\\
+So if the first phasis of "\backbone" task is sucessfully conduced, most of the
+projet delivrables will be carry through, even if some delivrable are lated
+or missing.
+\begin{figure}\leavevmode\center
+\begin{minipage}[t]{.4\linewidth}
+\includegraphics[width=1\linewidth]{dependence-dev}
+\caption{\label{dependence-dev}Dependence graph at development level}
+\end{minipage}\hfill\begin{minipage}[t]{.4\linewidth}
+\includegraphics[width=1\linewidth]{dependence-test}
+\caption{\label{dependence-test}Dependence graph at validation level}
+\end{minipage}
+\end{figure}
Index: /anr/section-4.2.tex
===================================================================
--- /anr/section-4.2.tex	(revision 12)
+++ /anr/section-4.2.tex	(revision 12)
@@ -0,0 +1,59 @@
+% FIXME EN FRANCAIS ET PAS A JOUR
+\mustbecompleted{
+Différents outils seront mis en place pour assurer le succès du projet, et
+faciliter la coopération entre les partenaires.
+\\
+La coordination générale du projet sera assurée par l'UBS/Lab-STICC. Le
+coordinateur assurera le management général du projet, il veillera au
+respect du plan d'exécution des travaux, et à la remise en temps voulu des
+fournitures. Il sera l'interlocuteur privilégié pour l'ANR, et représentera
+les autres partenaires du projet.
+\\
+La coordination technique sera assurée par l'UBS/Lab-STICC, qui sera chargé
+de garantir la cohérence scientifique et technique du projet au travers
+d'organisation des de Réunions Techniques d'Avancement, du suivi et de la
+centralisation des fournitures, etc.
+\\
+Des Réunions Techniques d'Avancement, regroupant tous les partenaires
+seront organisées une fois par trimestre. Ces réunions ont pour but de
+favoriser la circulation des informations techniques entre les différents
+partenaires. Des réunions de travail supplémentaires seront planifiées
+par chaque responsable de sous projet.
+\\
+Les outils logiciels développés par les partenaires académiques dans le
+cadre du projet XXXXX seront distribués en tant que logiciel libre. Ceci
+facilitera la circulation des informations techniques au sein mais aussi en
+dehors du projet.
+\\
+Pour faciliter la circulation des informations techniques et la coopération
+entre les différents partenaires, un serveur WEB accessible par Intranet
+sera mis en place par l'UBS/Lab-STICC dès le démarrage du projet. Ce
+serveur permettra à tous les partenaires d'accéder à la documentation
+technique partagée, ainsi qu'aux présentations effectuées à l'occasion des
+réunions techniques trimestrielles, etc.
+\\
+La validation des fournitures de ce projet sera assurée de façon continue
+par les auteurs de celles-ci ainsi que par les partenaires qui auront en
+charge soit l'intégration de ces fournitures soit produiront eux-même des
+éléments en relation avec la fourniture à valider.
+\\
+Le suivi du projet sera fait de la façon suivante: 
+\\
+Un compte-rendu d'activité sera communiqué tous les 6 mois au chargé du
+dossier par le coordinateur du projet. Ce rapport comprendra une
+contribution par partenaire et un document rédigé par le coordinateur.
+\\
+Un rapport sur chaque jalon sera fourni à T0+12, T0+24. Ce rapport donnera
+l'état d'avancement du projet avec de indicateurs tels que:
+1.respect des livraisons techniques (modèles et outils) \\
+2.conformité des livraisons aux spécifications \\
+3.dates de livraison des fournitures  \\
+4.état d'avancement du projet \\
+5.risques et actions correctives \\
+\\
+Un rapport final sera rendu au terme du projet a T0 + 36
+\\
+Une revue de projet annuelle, avec l'administration, sera organisée afin de
+faire un bilan détaillé de l'avancement du projet. Il pourra être décidé en
+commun des évolutions éventuelles du projet.
+}
Index: r/wp-IRISA.tex
===================================================================
--- /anr/wp-IRISA.tex	(revision 11)
+++ 	(revision )
@@ -1,79 +1,0 @@
-\documentclass[11pt,a4paper]{article}
-
-\usepackage[french]{babel}
-\usepackage[utf8x]{inputenc}
-\usepackage{times}
-\usepackage[T1]{fontenc}
-\usepackage{aeguill}
-\usepackage{verbatim}
-\usepackage{algorithm,algorithmic}
-\usepackage{xmpmulti}
-\usepackage{graphicx}
-\usepackage{color}
-
-\definecolor{gris25}{gray}{0.75}
-\definecolor{gris75}{gray}{0.30}
-
-
-\begin{document}
-
-\section {IRISA-WP}
-
-\subsection{Work Package 1 : Un compilateur reciblable pour MIPS Ã©tendu}
-
-DÃ©livrable : software
- 
-ImplÃ©mentation d'un back-end de compilation ciblant une version Ã©tendue du processeurs MIPS, et
-opÃ©rant Ã  partir de la reprÃ©sentation intermÃ©diaire commune dÃ©finie en \ref{?} et issue de GCC. 
-Ce \emph{back-end} intÃšgrera en particulier des passes d'extraction de motifs de calculs
-(sous-graphes), ainsi une passe de sÃ©lection d'intructions basÃ©e sur des techniques de couvertures 
-de graphes, permettant d'exploiter au mieux les motifs d'instruction ``spÃ©cialisÃ©s'' spÃ©cifiÃ©s 
-par l'utilisateur et/ou extraits Ã  partir de l'application.
-\textcolor{gris75}{Ici, il faut voir sir la RI prposÃ©e ne permettrait pas de regÃ©nÃ©rer un code 
-C dans lequel l'utilisation d'instruction spÃ©cialisÃ©e se fait au travers de directives de type
-\texttt{asm\{ \ldots \}} . une telle approche offrirait une flexibilitÃ© accrue, sans 
-impacter la qualitÃ©/performance des rÃ©sultats obtenus.}
-
-\subsection{Work Package 2 : DÃ©finition d'un modÃšle simplifiÃ© de micro-architecture MIPS extensible}
-
-DÃ©livrable : software
-
-DÃ©finition d'un modÃšle extensible de micro-architecture basÃ©e sur un processeur de type MIPS
-pipelinÃ© Ã  5 Ã©tage (incluant cache de donnÃ©es et instruction), offrant Ã  l'utilisateur 
-la possibilitÃ© de dÃ©finir ses propres extensions architecturales, au travers d'un Domain
-Specific Language (on exploitera les technologies d'IDM XText-EMF). 
-
-On mettra Ã©galement en {\oe}uvre un outil de gÃ©nÃ©ratiion de description matÃ©rielle synthÃ©tisable
-(VHDL) de la micro-architecture Ã  partir de ce modÃšle, en utilisant des techonologies d'IngÃ©nierie 
-dirigÃ©e par les modÃšles (EMF-XPAND)
- 
-Deux version de cet outil sont envisagÃ©es, dans la premiÃšre (qui est l'object de ce WP), on
-restreindra les possibilitÃ©s de communication entre le processeur et ses extension Ã  des communictaion passant par la file de
-regsitre du processeur (en permettant Ã©ventuellement un plus grand nombre d'accÃšs en
-lecture.Ã©criture par cycle).
-
-\subsection{Work Package 3 : DÃ©finition d'un modÃšle complexe de micro-architecture MIPS extensible}
-
-DÃ©livrable : rapport/software ?
-
-Dans la seconde version (plus orientÃ©e exploratoire) on souhaite pouvoir lever la limitaion portant
-sur les communications  et permettre un couplage plus fin entre les extensions et le coeur de la
-micro-rachgitcture, par exemple, en proposant un accÃšs direct au cache de donnÃ©es et/ou en donnant
-la possibiliteÃ© aux extensions de rÃ©utiliser les opÃ©rateurs (par exemple le multiplieur
-$32 \times 32$ \, bits) mis en oeuvre dans le chemin de donnÃ©e natif du processeur.
-
-Ici on pourrait Ã©galement envisager des connections directes avec d'autres composants au travers de
-structures similaire aux \emph{FSL} disponibles sur les processeurs softcore Microblaze de la
-sociÃ©tÃ© Xilinx.
-
-\subsection{Work Package 4 : DÃ©finition d'un modÃšle complexe de micro-architecture MIPS extensible}
-
-DÃ©livrable : rapport 
-
-Le dernier \emph{package} a Ã©galement un caractÃšre exploratoire, et portera sur l'intÃ©gration de ce
-type d'extension architecturales au sein d'un compilateur. En particulier, il s'agira d'Ã©tudier
-comment il est possible d'intÃ©grer ces instructions complexes dans la passe de selection de code,
-tout en s'assurant que leur contraintes d'utilisation soient respectÃ©es.
-
-
-\end{document}
